Clinical test scheme deviation automatic checking method and device and computer equipment

Automatically parsing PD records of CTMS and EDC systems through large language models, solving the problems of low verification efficiency and poor consistency in the existing technology, achieving efficient and accurate data consistency inspection and difference comparison, and improving the intelligence and compliance of clinical trial data management.

CN120257967AActive Publication Date: 2025-07-04HANGZHOU TIGERMED CONSULTING

Patent Information

Application Number
CN202510746150.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

When verifying the deviation of the scheme in CTMS and EDC systems, the prior art has problems such as low efficiency, easy human errors and difficult to achieve real-time monitoring. Especially in large-scale multi-center trials, data inconsistency and compliance issues between systems cannot be effectively solved.

Method used

The large language model is used to automatically analyze PD records of CTMS and EDC systems, and through data block classification, field mapping and unstructured description analysis, data consistency check and difference comparison are realized, and structured reports are generated.

Benefits of technology

Significantly improve verification efficiency and accuracy, reduce manual intervention, ensure system compatibility, and achieve seamless integration with existing workflows, improving the intelligence and compliance of data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257967A_ABST
    Figure CN120257967A_ABST
Patent Text Reader

Abstract

The invention discloses a clinical test scheme deviation automatic checking method and device and computer equipment. The method comprises the following steps: acquiring PD records and PD classification grading data from a CTMS system and an EDC system to obtain two pieces of data of the same project of the CTMS system and the EDC system; analyzing the two parts of data by using a large language model to obtain two parts of analysis results; performing consistency check on the two data according to the two analysis results, and comparing differences to obtain a detection result; generating a report containing consistency check and difference comparison according to a detection result; and outputting the report. By implementing the method provided by the invention, the checking efficiency, accuracy and system compatibility can be remarkably improved, the manual intervention and learning cost can be reduced, and seamless integration with the existing workflow can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data processing method, and more particularly to an automatic verification method, device and computer equipment for clinical trial protocol deviations. Background Art

[0002] In clinical trials, PD (Protocol Deviation) refers to the situation where the pre - determined clinical trial protocol is not followed during the implementation of the trial. For example, a subject misses a scheduled visit, has an incorrect dosing, or fails to complete certain tests as required. These deviations may be caused by researchers, subjects, or other relevant parties, and can affect the integrity of the trial, the reliability of the data, or the safety of the subjects. Therefore, accurately recording and reporting protocol deviations is crucial for ensuring the compliance of the trial, the integrity of the data, and the protection of the safety of the subjects.

[0003] In clinical trial management, CTMS (Clinical Trial Management System) and EDC (Electronic Data Capture) are core tools, and the protocol deviation data recorded by the two differ in source and function. First, the clinical trial management system is a software for managing the operation of clinical trials, focusing on the administrative and project management of the trial. Its main functions include trial planning and tracking, managing the trial schedule, milestones, and resource allocation; site management, tracking research site and investigator information; subject management, monitoring the progress of subject recruitment, screening, visits, etc.; document management, storing trial - related documents such as informed consent forms and ethical approvals; and finance and compliance, managing budgets and payments, and ensuring compliance with regulatory requirements. CTMS can provide a global perspective to help the research team optimize operations and ensure that the trial progresses as planned. The electronic data capture system, on the other hand, is used to collect and manage clinical trial data, replacing traditional paper - based data collection. Its core functions include directly collecting clinical data of subjects, such as medical history and laboratory results; built - in data verification rules to ensure data accuracy; real - time data monitoring to reduce data entry errors; and providing data security guarantees to ensure compliance and provide an audit trail function. EDC aims to improve data quality and collection efficiency, ensuring data reliability and traceability.

[0004] The CTMS and EDC systems have different functional focuses, but there is an intersection between them. CTMS pays more attention to the management of trial operations (such as site information and subject status), while EDC focuses on the collection and verification of clinical data. Due to the different functions and data entry methods of the two systems, PD records may be inconsistent, missing, or misclassified. These inconsistencies may lead to data loss or errors, thus affecting the reliability of trial results and the validity of statistical analysis. In addition, inconsistent records may cause compliance issues and increase regulatory risks. Therefore, it is crucial to verify the differences in protocol deviation records between the CTMS and EDC systems.

[0005] Currently, the methods for verifying protocol deviations in clinical trials can be roughly divided into two types: manual verification and semi-automated verification. First, manual verification is a traditional verification method. The method is to export data related to protocol deviations from the CTMS and EDC systems respectively, and then manually compare the records by data managers to check for differences, omissions, or duplicates. The differences will be sorted into a report and the problem types and possible causes will be noted. However, the drawback of this method is that it has a large workload, especially in large multi-center trials, with low efficiency and is prone to human errors, such as missed checks or misjudgments. In addition, it is difficult to achieve real-time monitoring with manual verification, which may affect the progress of the trial. Semi-automated verification, on the other hand, uses technical tools to extract data from the CTMS and EDC systems and convert the data into a unified format. Preliminary screening is carried out through simple rules to identify differences. This method is more efficient than manual review and can reduce repetitive work, but still requires certain technical capabilities, and the writing and maintenance of scripts will bring additional work. For descriptive texts or complex protocol deviations, it is difficult to handle automatically, so it cannot completely replace manual review.

[0006] Therefore, it is necessary to design a new method that significantly improves verification efficiency, accuracy, and system compatibility, reduces manual intervention and learning costs, and ensures seamless integration with the existing workflow. Summary of the Invention

[0007] The purpose of the present invention is to overcome the defects of the prior art and provide an automatic verification method, device, and computer device for clinical trial protocol deviations.

[0008] To achieve the above purpose, the present invention adopts the following technical solutions: The automatic verification method for clinical trial protocol deviations includes: Obtaining PD records and PD classification and grading data from the CTMS system and the EDC system to obtain two sets of data for the same project in the CTMS system and the EDC system; Using a large language model to parse the two sets of data to obtain two parsing results; Perform consistency checks on the two pieces of data based on the two parsing results, and compare the differences to obtain a detection result; Generate a report including consistency verification and difference comparison based on the detection result; Output the report.

[0009] A further technical solution thereof is: parsing the two pieces of data by using a large language model to obtain two parsing results, including: Classify the two pieces of data by using a large language model to obtain deviation record data and classification and grading data; Automatically map different system fields to the deviation record data and the classification and grading data and parse the unstructured PD description to extract key information; Convert the key information into a unified internal data structure to obtain two parsing results.

[0010] A further technical solution thereof is: the classifying the two pieces of data by using a large language model to obtain deviation record data and classification and grading data, including: Split the two pieces of data into data blocks and classify each data block by using the zero-shot classification ability of the large language model, and combine a voting mechanism to determine the final category to obtain deviation record data and classification and grading data.

[0011] A further technical solution thereof is: the automatically mapping different system fields to the deviation record data and the classification and grading data and parsing the unstructured PD description to extract key information, including: Construct a prompt; Input the deviation record data and the classification and grading data combined with the prompt into the large language model to parse the unstructured PD description to obtain a parsing result in JSON format; Deserialize the parsing result in JSON format into an object in a programming language to obtain key information.

[0012] A further technical solution thereof is: the performing consistency checks on the two pieces of data based on the two parsing results, and comparing the differences to obtain a detection result, including: Check the internal consistency of the two parsing results through the large language model to obtain a consistency result; Compare the two parsing results by subject and time, mark the missing records and evaluate the consistency of the two pieces of deviation record data to determine the missing marks and reasons for differences; Wherein, the detection result includes a consistency result, missing marks and reasons for differences.

[0013] Its further technical solution is: checking the internal consistency of the two parsing results through a large language model to obtain a consistency result, including: Constructing a prompt word according to the two parsing results; Inputting the two parsing results and the prompt word into the large language model for internal consistency analysis to obtain a consistency result; Among them, the consistency result includes a consistency result, a confidence level, a reason for inconsistency, and a suggestion.

[0014] Its further technical solution is: comparing the two parsing results by subject and time, marking missing records and evaluating the consistency of two deviation record data to determine the missing marks and reasons for differences, including: Grouping the two parsing results according to the subject screening number, and classifying the records under the same subject according to the PD type to obtain multiple groups of data; Sorting each group of data according to the event occurrence time to obtain a sorting result; Constructing a comparison prompt word; Inputting the sorting result and the comparison prompt word into the large language model to obtain a comparison result, where the comparison result includes the situation of marked missing and the reason for the difference.

[0015] Its further technical solution is: generating a report including consistency verification and difference comparison according to the detection result, including: Generating a structured Excel file including consistency verification and difference comparison according to the detection result to obtain a report.

[0016] The present invention also provides an automatic verification device for clinical trial protocol deviations, including: An acquisition unit, configured to acquire PD records and PD classification and grading data from a CTMS system and an EDC system to obtain two sets of data of the same project in the CTMS system and the EDC system; An analysis unit, configured to analyze the two sets of data by using a large language model to obtain two parsing results; A PD inspection unit, configured to perform consistency inspection on the two sets of data according to the two parsing results and compare the differences to obtain a detection result; A report generation unit, configured to generate a report including consistency verification and difference comparison according to the detection result; An output unit, configured to output the report.

[0017] The present invention also provides a computer device, the computer device includes a memory and a processor, a computer program is stored on the memory, and when the processor executes the computer program, the above method is implemented.

[0018] The beneficial effects of the present invention compared with the prior art are as follows: By obtaining PD records and their classification and grading data from the CTMS system and the EDC system, and using the large language model to parse the data, the system can efficiently and automatically process the consistency check and difference comparison of the data; this process significantly improves the efficiency and accuracy of the verification, ensures the compatibility between different systems, and reduces the manual intervention and learning cost. By generating reports containing consistency verification and difference comparison, the system can seamlessly integrate the existing work processes, realize the automation and optimization of the processes, and greatly improve the work efficiency and quality.

[0019] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic diagram of the application scenario of the automatic verification method for clinical trial protocol deviations provided by the embodiments of the present invention; Figure 2 It is a schematic flowchart of the automatic verification method for clinical trial protocol deviations provided by the embodiments of the present invention; Figure 3 It is a schematic sub - flowchart of the automatic verification method for clinical trial protocol deviations provided by the embodiments of the present invention; Figure 4 It is a schematic sub - flowchart of the automatic verification method for clinical trial protocol deviations provided by the embodiments of the present invention; Figure 5 It is a schematic sub - flowchart of the automatic verification method for clinical trial protocol deviations provided by the embodiments of the present invention; Figure 6 It is a schematic sub - flowchart of the automatic verification method for clinical trial protocol deviations provided by the embodiments of the present invention; Figure 7 It is a schematic sub - flowchart of the automatic verification method for clinical trial protocol deviations provided by the embodiments of the present invention; Figure 8 It is a schematic block diagram of the automatic verification device for clinical trial protocol deviations provided by the embodiments of the present invention; Figure 9 It is a schematic block diagram of the computer device provided by the embodiments of the present invention. Detailed Embodiments

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0024] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0025] It should be further understood that the term " / and" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0026] Please refer to Figure 1 and Figure 2 , Figure 1 which is a schematic diagram of the application scenario of the automatic verification method for deviations in the clinical trial protocol provided by the embodiments of the present invention. Figure 2 which is a schematic flowchart of the automatic verification method for deviations in the clinical trial protocol provided by the embodiments of the present invention. This automatic verification method for deviations in the clinical trial protocol is applied to a server. The server interacts with the terminal, and through the application of the large language model, the verification efficiency, accuracy, and system compatibility are significantly improved. This method reduces manual intervention by automatically parsing the data from the CTMS and EDC systems for consistency checking and difference comparison; it effectively improves the precision of data processing and ensures the consistency between data by using the large language model for classification, automatic mapping of fields, and parsing of unstructured data; in addition, this method can automatically generate reports and export them in a structured format, simplifying the report generation process, reducing the learning cost, ensuring seamless integration with the existing workflow, and greatly improving the automation and intelligence level of verification.

[0027] Figure 2 is a flowchart of the automatic verification method for deviations in the clinical trial protocol provided by the embodiments of the present invention. As Figure 2 shown, this method includes the following steps S110 to S150.

[0028] S110. Obtain PD records and PD classification and grading data from the CTMS system and the EDC system to obtain two sets of data for the same project in the CTMS system and the EDC system.

[0029] In this embodiment, the two sets of data for the same project in the CTMS system and the EDC system are the PD records and PD classification and grading data of the CTMS system and the PD records and PD classification and grading data of the EDC system.

[0030] Obtain data related to deviations from the clinical trial management system and the electronic data capture system.

[0031] CTMS is a system for managing and monitoring clinical trials. Here, the CTMS system records the visit data of the subjects and the deviations during the trial process.

[0032] Through the API interface or file export method, the system extracts data from the CTMS system. The extracted data includes: Structured fields: such as subject ID (subject_id), visit date (visit_date), deviation type (deviation_type), etc.

[0033] Unstructured data: For example, the description text of PD (such as "The subject missed the 3rd visit due to traffic problems").

[0034] EDC is a system for recording and managing data in clinical trials, usually used to collect patient data and various records during the trial process.

[0035] Similar to the CTMS system, the EDC system also provides an API interface or supports the file export function to obtain data related to PD, and the content and structure are usually different from the data format in the CTMS.

[0036] The data extracted from the EDC system also includes structured fields (such as subject ID, visit date) and unstructured description text (such as PD description text).

[0037] The types of data obtained include: PD records: Record the deviations or events of the subjects during the clinical trial process.

[0038] PD classification and grading data: Refers to the classification of PD records (such as "visit delay", "medication error", etc.) and grading information (such as "minor", "major", etc.).

[0039] To access data through the interface between CTMS and EDC, login and authentication are usually required. Through the interface, the system can automatically pull project data, including PD records and classification information.

[0040] If the API interface cannot be used, data can also be obtained by importing files in formats such as CSV and Excel. In this case, the file format needs to conform to the specified structure to ensure correct data parsing. There is no need to adapt to different CTMS and EDC systems.

[0041] In the whole system, this step is very crucial because it is the basis for subsequent data parsing and verification. Any issues with data acquisition and quality will directly affect the results of subsequent analysis and verification.

[0042] The accuracy and integrity during data acquisition ensure the successful execution of subsequent steps (such as data parsing, verification, etc.).

[0043] Step S110 is the first link in the entire automated verification system, aiming to obtain two sets of data on PD records through the API interface of the CTMS and EDC systems or file import. These data provide the basis for subsequent intelligent parsing and verification modules, ensuring data consistency and accuracy.

[0044] In this embodiment, PD record data is sensitive privacy data within the enterprise. For data security considerations, this data is not allowed to be transmitted to the cloud. Therefore, it is not possible to use large models deployed externally in the cloud for data processing. All data processing must be completed on the enterprise's internal servers or workstations to reduce the risk of data leakage. Therefore, the enterprise needs to privately deploy large models in its internal environment.

[0045] Ollama is an open-source tool designed to efficiently run large language models in a local environment. It helps enterprises achieve secure, efficient, and low-cost AI deployment. In combination with Nginx reverse proxy, Ollama can ensure that the interfaces of large models are not accessible without authorization. The following is the specific process for deploying Ollama and Nginx reverse proxy: First, it is necessary to evaluate the existing hardware environment to ensure that it can support the implementation of subsequent steps. This evaluation process includes checking whether key resources such as the processor, memory, storage space, and network connection are sufficient, and adjusting or upgrading the hardware if necessary to meet the deployment requirements.

[0046] After confirming that the hardware environment meets the requirements, the next step is to download and install the "Ollama" software. This process usually involves accessing the official website or the specified download link to obtain the latest installation package and completing the software installation according to the installation instructions.

[0047] After installing Ollama, the next step is to download and run the "Qwen2.5 32B" large model. This usually means obtaining the pre-trained model file from a specific model library or platform, and then using Ollama or other relevant tools to load and start the model for subsequent applications or experiments.

[0048] To optimize the system's network communication and improve its security and stability, the next step is to install and configure Nginx as a reverse proxy server. This process includes downloading the Nginx installation package, executing the installation command, and modifying the Nginx configuration file according to actual needs to achieve functions such as load balancing and SSL encryption, ensuring the efficient operation and stability of the system.

[0049] Finally, to further enhance the system's security, an HTTP basic authentication mechanism needs to be configured for Nginx. This configuration can be achieved by editing the Nginx configuration file and adding the corresponding authentication instructions and user credentials. After completing these settings, only authenticated clients can access the protected resources, thereby improving the system's security protection ability.

[0050] Through the above steps, enterprises can ensure the efficient and secure deployment and operation of large language models in a local privatized environment, effectively avoid the risk of data leakage, and ensure the control and security of system access.

[0051] S120. Use a large language model to parse the two pieces of data to obtain two parsing results.

[0052] In this embodiment, the two parsing results refer to the key information corresponding to the PD record data and the PD classification and grading data obtained respectively through classification and parsing.

[0053] In one embodiment, please refer to Figure 3 , the above step S120 may include steps S121 to S123.

[0054] S121. Use a large language model to classify the two pieces of data to obtain deviation record data and classification and grading data.

[0055] In this embodiment, the two pieces of data are split into data blocks, and the zero-shot classification ability of the large language model is used to classify each data block, and the voting mechanism is combined to determine the final category to obtain deviation record data and classification and grading data.

[0056] Split the two input data (usually from different CTMS and EDC systems), and use the zero-shot classification ability of the large language model to classify each data block. In this way, the model can identify the types of data blocks, such as PD record data and PD classification and grading data.

[0057] First, split the two data into multiple data chunks. Each data chunk contains several rows of data, and the specific number is usually determined according to the input limit of the model (for example, 5 to 10 rows). Each data chunk is processed as an independent input item through the large language model.

[0058] Next, use the large language model to classify each data chunk. Due to the zero-shot learning ability of the large language model, the system can identify the category of the data chunk according to the prompt without explicit training. For example, the data chunk may be labeled as "PD record", "PD classification and grading", or "other".

[0059] To improve the classification accuracy, a voting mechanism is adopted. Specifically, the classification results of multiple data chunks are aggregated through the inference ability of the large language model to determine the final category of each data chunk. Each data chunk will return a classification result and its confidence level, and the final category will be determined according to the aggregated result of the confidence levels.

[0060] Through these steps, the system can automatically identify and classify PD record data and classification and grading data, eliminating the need for manual classification or traditional hard-coded classification.

[0061] S122. Automatically map different system fields to the deviation record data and classification and grading data and parse the unstructured PD description to extract key information.

[0062] In one embodiment, refer to Figure 4 , the above step S122 may include steps S1221 to S1223.

[0063] S1221. Construct a prompt. S1222. Input the deviation record data and classification and grading data combined with the prompt into the large language model to parse the unstructured PD description and obtain a parsed result in JSON format. S1223. Deserialize the JSON format parsed result into an object of a programming language to obtain key information.

[0064] In this embodiment, the key information refers to the key fields extracted from the PD record data and PD classification and grading data, such as subject ID, deviation type, classification code, etc.

[0065] In this embodiment, the data fields obtained from different CTMS and EDC systems are automatically mapped to a unified data structure, and the unstructured PD descriptions are parsed to extract key information. This process aims to address the differences in field names and data formats between different systems and ensure data consistency.

[0066] Since the field names in CTMS and EDC systems may vary (e.g., "subject_id" vs. "patient_id"), it is necessary to automatically map the field names in different systems. Through its semantic understanding ability, the large language model can identify these differences and perform automatic mapping. For example, mapping "patient_id" to "subject_id" to ensure that data from different systems can be uniformly understood.

[0067] In addition to structured fields, PD records may also contain unstructured descriptions (e.g., "The subject missed the 3rd visit due to transportation issues"). Using its semantic understanding ability, the large language model can extract key information from these texts, such as deviation types, deviation reasons, etc. This step is crucial for processing complex and irregular text data.

[0068] Through these mapping and parsing processes, the system can ensure that data from different systems is standardized and structured, facilitating subsequent processing and analysis.

[0069] S123. Convert the key information into a unified internal data structure to obtain two parsing results.

[0070] In this embodiment, the classified and parsed key information is converted into a unified internal data structure for subsequent data processing, verification, and analysis.

[0071] In this embodiment, two main data structures are defined: PD record data structure (PD_RECORD), including fields such as "subject_id", visit date, category, description text, level, etc.

[0072] PD classification and grading data structure (PD_TYPE_LEVEL), including fields such as main classification code, sub-classification code, classification name, and grading.

[0073] Through the design of prompting words of the large language model, the extracted key information is converted into a unified data structure. This process will automatically select the corresponding structure according to the type of input data (PD record or classification and grading data), and fill the parsed field values into the structure. The final structured data can be represented in JSON format to ensure data standardization.

[0074] After completing this step, the system will output two structured data results, which will facilitate consistency checking and comparative analysis in the subsequent PD verification module.

[0075] Step S120 solves the problem of data format differences between the CTMS and EDC systems through a series of operations, using the powerful semantic understanding and adaptive capabilities of the large language model, and automatically classifies, parses and structures the data from different systems. Specifically, step S121 ensures that the two sets of data can be effectively parsed and unified by segmenting and classifying the data, S122 automatically maps and parses unstructured descriptions through field mapping, and S123 converts key information into a unified data structure. Through this process, the system achieves efficient and automated PD record verification and data integration, significantly improving the efficiency and accuracy of data processing.

[0076] Specifically, through the semantic understanding ability of the large language model (LLM), the intelligent parsing and structured processing of patient data (PD) records and classified and graded data in the clinical trial management system (CTMS) and electronic data capture system (EDC) is realized. This module breaks through the limitations of traditional methods in processing heterogeneous data and significantly improves the automation level, adaptability and maintainability of the system.

[0077] In clinical trial data management, CTMS and EDC systems usually store PD records and classification and grading data in different formats. Although these data are highly regular (i.e., the field structure is relatively standardized), there are still some subtle format differences, such as different field names (e.g., "subject_id" vs. "patient_id"), inconsistent date formats (e.g., "YYYY-MM-DD" vs. "DD / MM / YYYY"), and differences in classification and grading expressions (e.g., "minor deviation" vs. "minor deviation"). Traditional methods usually adapt to different CTMS and EDC systems one by one through hard coding, requiring specific parsing logic to be written for each system, mapping field names, and converting formats, which increases the amount of code and complexity. When the CTMS or EDC system is updated (e.g., API changes or field adjustments), the code needs to be modified and the system redeployed, which is time-consuming and labor-intensive, and has poor scalability. Therefore, when a new system is added, a new adapter must be developed, which cannot quickly adapt to diverse data sources.

[0078] In order to solve these problems, this embodiment uses the semantic understanding ability and adaptability of LLM to automatically parse and structure PD records and classified and graded data, fundamentally solving the limitations of traditional methods.

[0079] Specifically, PD records (including structured fields and unstructured descriptions) and classification and hierarchical data are identified from the input data.

[0080] Utilize the semantic understanding ability of the LLM to automatically adapt to the heterogeneous data formats of different CTMS and EDC systems without hard - coding adaptation.

[0081] Identify field semantics and automatically map field names of different systems (e.g., map "patient_id" to "subject_id"), including the automatic mapping of Chinese and English field names.

[0082] Parse unstructured PD descriptions and extract key information (e.g., deviation type, reason).

[0083] Guide the LLM through prompts (Prompt) to convert the parsed data into a unified internal data structure.

[0084] Automatic search locates PD classification and grading data and PD record data from the two provided data sets. Formally, automatic search can be regarded as a text classification process: split the input data table into appropriately sized data chunks (chunk), use the zero - shot classification ability of the large - language model to classify each data chunk, and determine the final category of the data chunk through a voting mechanism, thus achieving the location of the required data. The following describes it in detail from four aspects: data pre - processing, classification process, voting mechanism, and result output.

[0085] Assume that the input data can be divided into K data tables, and assume that one of the data tables is D, which contains N rows of data: D = {d1, d2,..., dN}; To improve classification efficiency and model understanding ability, combine consecutive (k) rows of data into a data chunk (chunk) to form a data chunk set (C).

[0086] C = {c1, c2,..., cM}, ; where cj is the (j) - th data chunk, which contains (k) rows of data (the last data chunk may have less than (k) rows). k is usually selected according to the data scale and model input limitations (such as the maximum number of tokens), and typical values are 5 - 10 rows.

[0087] Utilize the zero - shot classification ability of the large - language model to classify each data chunk cj to determine whether it contains PD classification and grading data or PD record data. The classification label set L = {"PD record", "PD classification and grading", "other"} represents the three possible categories of the data chunk. Through prompt design, input each data chunk into the large - language model for classification. The model returns the classification result (e.g., "PD record") and gives the confidence level P.

[0088] Due to the volatility of the results returned by the large model, a single classification result may not be accurate. Therefore, the classification results are aggregated through a voting mechanism to determine the final category of the data block. For the data block set C of each data table D, the results (Lj, Pj) of each data block Cj in C are obtained, and the sum of the confidence levels for each category L is calculated. The category L with the largest sum of confidence levels is selected as the category of the data table.

[0089] The PD classification and grading data structure is designed as follows: json { "main_category_code": "string", / / Main classification code; "main_category_name": "string", / / Main classification name "sub_category_code": "string", / / Sub - classification code; "sub_category_name": "string", / / Sub - classification name; "level": "string" / / Grading; } The PD record data structure is designed as follows: json { "subject_id": "string", / / Subject ID; "visit_date": "string", / / Visit date (format: YYYY - MM - DD); "main_category_name": "string", / / Main classification name; "sub_category_name": "string", / / Sub - classification name; "description": "string", / / PD description text (such as "The subject missed the 3rd visit due to transportation problems"); "level": "string", / / Classification and grading; "source": "string" / / Data source ("CTMS" or "EDC"); } After determining the category of the data table (such as PD records or classified and graded data) through automatic search and defining a unified data structure, the data parsing module utilizes the semantic understanding ability of the large language model (LLM) to automatically parse the content in the data table and map the required fields to the unified data structure. This process uses JSON as an intermediary to ensure the structuring and consistency of data during parsing and mapping. Specifically as follows: Construct a prompt and input it into the large language model. The prompt is constructed as follows: The text of the data row is {text}; It comes from the {type} data table; Please try to parse it into an array of JSON objects with the {struct} structure. If a certain row cannot be parsed, output a JSON object with the string UNKOWN Where text is the text of the data row, type is the type of the data table (which may be a PD classification and grading table or a PD record table), and struct is the structure determined according to type. If type is a PD classification and grading table, then struct is the PD_TYPE_LEVEL structure; if type is a PD record table, then struct is the PD_RECORD structure.

[0090] Deserialize the JSON string returned by the large language model into an object in the programming language to complete data parsing.

[0091] S130. Perform a consistency check on the two pieces of data according to the two parsing results, and compare the differences to obtain a detection result.

[0092] In this embodiment, the detection result refers to the final output generated based on the consistency check and data comparison, including the comparison result of PD records in the two systems, the marking of missing records, and the reasons for the differences found during the comparison process. The detection result provides a basis for data cleaning and quality control, helping to identify and correct inconsistent or missing problems in the data.

[0093] The process of performing a consistency check and comparison verification on PD (adverse event) records in CTMS (Clinical Trial Management System) and EDC (Electronic Data Capture System) through the large language model (LLM). Specifically, utilize the semantic understanding and reasoning ability of the LLM to check the consistency of the data, identify missing data, and analyze the differences between the two systems.

[0094] This step is responsible for performing a consistency check and difference comparison on the two data parsing results extracted from CTMS and EDC, and finally obtaining a detection result. The detection result includes the consistency evaluation of the data, the marking of missing records, and the analysis of the reasons for the differences.

[0095] In one embodiment, please refer to Figure 5 , step S130 described above may include steps S131 to S132.

[0096] S131. Check the internal consistency of the two parsing results through a large language model to obtain a consistency result.

[0097] In this embodiment, the consistency result refers to an evaluation result obtained by performing semantic analysis on the PD records of the same subject and the same visit in the CTMS and EDC systems through a large language model, and determining whether elements such as data description, category, and time in the two systems match. If the data in the two systems is consistent at the semantic level, it is judged as consistent; otherwise, it is inconsistent, and the specific reason for the inconsistency is given.

[0098] In one embodiment, please refer to Figure 6 , step S131 described above may include steps S1311 to S1312.

[0099] S1311. Construct a prompt based on the two parsing results; S1312. Input the two parsing results and the prompt into the large language model for internal consistency analysis to obtain a consistency result; wherein, the consistency result includes a consistency result, a confidence level, a reason for inconsistency, and a suggestion.

[0100] In this embodiment, first, an internal consistency check is performed on the two parsing results. Specifically, it includes the following sub-steps: Based on the data in the CTMS and EDC systems, construct a prompt (Prompt). The prompt will include information such as "description", "main_category_name", "sub_category_name", etc., with the purpose of providing it to the LLM for further analysis.

[0101] Input the constructed prompt and the parsing results of CTMS and EDC into the large language model for consistency check. The large language model automatically evaluates whether the description and the category are consistent according to its deep semantic understanding ability. For example, the model will check whether there is a semantic logic match between the "visit delay" category and "the subject missed the 3rd visit due to traffic problems".

[0102] The results include the following aspects: Consistency result (whether it is consistent); Confidence level (between 0 and 1, indicating the reliability of the consistency); Reason for inconsistency (if there is an inconsistency, the model will infer the reason and give a suggestion).

[0103] S132. Compare the two parsed results by subject and time, mark the missing records, and evaluate the consistency of the two sets of deviation record data to determine the missing marks and reasons for differences; Among them, the detection results include the consistency results, missing marks, and reasons for differences.

[0104] In one embodiment, please refer to Figure 7 , the above step S132 may include steps S1321 to S1324.

[0105] S1321. Group the two parsed results according to the subject screening number, and classify the records under the same subject according to the PD type to obtain multiple groups of data; S1322. Sort each group of data according to the event occurrence time to obtain a sorted result; S1323. Construct a comparison prompt; S1324. Input the sorted result and the comparison prompt into a large language model to obtain a comparison result, where the comparison result includes the missing mark situation and the reason for the difference.

[0106] In this embodiment, step S132 mainly focuses on comparing the PD records of the same subject and the same visit in the CTMS and EDC systems to identify data differences and omissions. This process includes the following sub-steps: Group the PD records in CTMS and EDC according to the subject's screening number (such as a unique ID). Ensure that each group represents the data of the same subject.

[0107] Among the data of the same subject, further classify according to the PD type. This is done to ensure a unified data structure for comparison and avoid incorrect comparison due to different record categories.

[0108] Sort each group of data according to the event occurrence time. By sorting in chronological order, it can ensure that the same event (such as a delay in a certain visit) in the two systems is correctly compared.

[0109] Construct a comparison prompt. The content of the prompt includes the list of PD records of the same subject and the same visit in the CTMS and EDC systems, and requires the LLM to compare according to elements such as description and time to identify the differences between the records. The comparison results are divided into three categories: Both systems have records and are consistent; CTMS has a record but EDC does not (marked as "missing"); EDC has a record but CTMS does not (marked as "missing").

[0110] The output of the LLM will include cases of marked missing records and an analysis of the reasons for the differences (such as data entry errors or system synchronization issues).

[0111] In this embodiment, the final detection results include the following key elements: Consistency result: Through internal consistency checks, evaluate whether the two data sets match semantically.

[0112] Missing markers: Compare the PD records in the two systems and mark which records are missing in one system.

[0113] Reasons for differences: Analyze and explain the reasons for data inconsistencies between the two systems, such as data errors, system synchronization issues, etc.

[0114] This embodiment realizes a comprehensive automated verification of PD records in CTMS and EDC systems by making full use of the powerful semantic analysis and reasoning capabilities of large language models. Through the following technical measures: Intelligent prompt generation: Construct prompts based on actual data and pass them to the large language model for verification; Automated consistency checking and comparison verification: Without manual intervention, it can perform efficient consistency checking and difference analysis on massive data; Deep reasoning and corrective suggestions: Provide reasonable analysis and suggestions for inconsistent records to support data cleaning.

[0115] This step not only improves the verification efficiency and accuracy, but also reduces manual intervention, significantly enhancing the intelligence and compliance of data management.

[0116] In the method of this embodiment, each PD record in the CTMS and EDC systems is verified to check its internal consistency and ensure the semantic consistency between the PD category (such as "visit delay") and the PD description (such as "the subject missed the 3rd visit due to traffic problems"). With the zero-shot text classification ability of the large language model (LLM), inconsistent records are intelligently identified and marked, and corrective suggestions are provided. This provides necessary support for subsequent data cleaning and quality control.

[0117] Compare the PD records of the same subject and the same visit in the CTMS and EDC systems to identify the differences between the two systems (such as missing records, classification discrepancies, description inconsistencies, etc.). With the deep reasoning ability of the LLM, the module can not only discover the differences, but also infer their root causes (such as data entry errors or system synchronization issues) and evaluate the severity of the differences, thus helping to improve data consistency.

[0118] The method of this embodiment gives full play to the advantages of large language models in semantic analysis and reasoning, realizes a fully automated process from consistency checking to difference comparison, significantly improves the verification efficiency and accuracy, and reduces the need for manual intervention. In addition, the module design has flexible scalability, can adapt to clinical trials of different scales and diverse CTMS / EDC systems, and provides intelligent data management and regulatory compliance solutions.

[0119] Specifically, traverse each PD record in CTMS and EDC, and extract its main_category_name, sub_category_name, and description fields.

[0120] Construct a prompt, referring to the following structure: Description: {description}; Category: {main_category_name}, {sub_category_name}; Is the description consistent with the category? Please return the consistency result (consistent / inconsistent), confidence level, and the reason and suggestion when inconsistent.

[0121] The result returned by the LLM is in text format. It is necessary to parse the returned text and extract the consistency result ("consistent" or "inconsistent"), confidence level (0 - 1), reason for inconsistency, and suggestion. To improve efficiency, batch prompts can be constructed.

[0122] The comparison and verification process of PD records between CTMS and EDC is as follows: Group the PD records in CTMS and EDC according to the subject screening number.

[0123] For the PD records under each subject screening number, group them according to the PD type.

[0124] For the PD records of the same type under the same subject screening number, sort them in chronological order.

[0125] Construct a prompt, referring to the following structure: The list of PD records from CTMS is: {CTMS_PD_RECORDS}; The list of PD records from EDC is: {EDC_PD_RECORDS}; Please compare one - by - one the records in the list of PD records from CTMS and the list of PD records from EDC, and infer whether they are the same PD according to the description and time. The comparison results are divided into three categories: There are records in both systems and they are consistent.

[0126] CTMS has it, EDC doesn't: Marked as "missing".

[0127] EDC exists, CTMS does not: Marked as "missing".

[0128] The LLM returns the result in text format. It is necessary to parse the returned text and generate the required data structure object according to the parsing result.

[0129] The prompt words need to be optimized according to the specific situation.

[0130] Through the above technical implementation steps, the data verification work can be completed efficiently and accurately, further improving the quality and compliance of clinical trial data management.

[0131] S140. Generate a report containing consistency verification and difference comparison based on the detection result.

[0132] In this embodiment, the report refers to a structured Excel file generated based on the detection result, including the consistency verification of PD records and the difference comparison between the CTMS and EDC systems, aiming to provide detailed verification results and corrective suggestions.

[0133] Specifically, generate a structured Excel file containing consistency verification and difference comparison based on the detection result to obtain the report.

[0134] The result recording module is the core output part of the PD record verification system, responsible for saving and displaying the verification result in the form of a structured Excel file, facilitating users to view, analyze and archive. The main functions of this module are as follows: Record the verification result: Verification result of the CTMS system: Record the internal consistency verification result of PD records in the CTMS system, and check whether the PD category matches the description.

[0135] Verification result of the EDC system: Record the internal consistency verification result of PD records in the EDC system, and also check the consistency of the PD category and the description.

[0136] Verification result of the comparison between the CTMS and EDC systems: Record the verification result of the comparison of PD records in the CTMS and EDC systems, and identify the differences between the two.

[0137] Generate a structured report: Output the verification result in the form of an Excel file, divided into three independent worksheets, corresponding to the above three parts respectively.

[0138] Provide a clear field structure to facilitate users to quickly understand the verification result.

[0139] The Excel file generated by the result recording module contains three worksheets, each corresponding to the verification results of a part, with a clear structure and reasonable field design. The following is the detailed structure of each part: Overall structure of the Excel file File name: PD_Verification_Result_[Timestamp].xls; Worksheets: CTMS_Consistency: Verification results of the internal consistency of CTMS PD records.

[0140] EDC_Consistency: Verification results of the internal consistency of EDC PD records.

[0141] CTMS_vs_EDC_Differences: Verification results of the comparison between CTMS and EDC PD records.

[0142] Field structure of the CTMS_Consistency worksheet; Record_ID: The unique identifier of the PD record (composed of the subject ID and visit number).

[0143] Visit_Date: Visit date.

[0144] Deviation_Type: PD category.

[0145] Description: PD description text.

[0146] Consistency_Result: Consistency check result (marked as "consistent" or "inconsistent").

[0147] Confidence_Score: Confidence level of the consistency check (range 0 - 1, based on zero-shot classification of large language models).

[0148] Inconsistency_Details: If there is an inconsistency, provide detailed description (e.g., "The description is visit delay, but the category is marked as administration error").

[0149] Suggestion: Corrective suggestion (e.g., "Suggest updating the category to 'visit delay'").

[0150] EDC_Consistency worksheet: The field structure of this worksheet is the same as that of CTMS_Consistency, but the data is sourced from the EDC system.

[0151] CTMS_vs_EDC_Differences Worksheet: This worksheet is grouped by subject ID for easy tracking of subjects.

[0152] Record_ID: The unique identifier of the PD record.

[0153] Visit_Date: The visit date (the dates in CTMS and EDC may not be the same).

[0154] Deviation_Type_CTMS: The PD category in CTMS.

[0155] Deviation_Type_EDC: The PD category in EDC.

[0156] Description_CTMS: The PD description in CTMS.

[0157] Description_EDC: The PD description in EDC.

[0158] Difference_Type: The type of difference (e.g., "Missing", "Value mismatch", "Semantic inconsistency").

[0159] Missing: One party has a record while the other does not.

[0160] Value mismatch: There is an inconsistency between structured fields (such as categories).

[0161] Semantic inconsistency: There are differences in the description semantics.

[0162] Similarity_Score: The semantic similarity of the description (0 - 1, based on semantic vector comparison of large language models).

[0163] Severity: The severity of the difference (e.g., "Minor" or "Major").

[0164] Details: The detailed information of the difference (e.g., "The CTMS record exists, but the EDC record is missing").

[0165] Suggestion: Suggestions for correction (e.g., "It is recommended to supplement the record in EDC").

[0166] S150. Output the said report.

[0167] In this embodiment, the said report is output to the terminal.

[0168] The method of this embodiment obtains PD (Product Data) records by interfacing with CTMS (Clinical Trial Management System) and EDC (Electronic Data Capture) systems or through file import. Utilizing the semantic understanding ability of large language models, it intelligently parses the obtained PD records, automatically adapts to the heterogeneous data formats and description methods between different systems, converts the originally unstructured text data into structured data, and unifies the format to provide consistent input data for subsequent verification. Based on the parsed data, it performs consistency verification of PD records and compares the PD records in CTMS and EDC systems. By precisely matching structured fields and semantically analyzing unstructured descriptions, it identifies inconsistent records (such as missing, misclassified, or semantically inconsistent ones), and classifies them according to the type and priority of the differences to ensure the accuracy and operability of the verification results. It generates a detailed verification report, clearly showing the difference details, semantic analysis results, and providing corresponding corrective suggestions.

[0169] Through development based on the Microsoft Office ecosystem, it can run in an environment familiar to users, reduce operation barriers, and achieve seamless integration with existing workflows, enhancing compatibility and convenience. By using the semantic understanding ability of large language models, it can automatically adapt to the PD record formats in different CTMS and EDC systems in the data parsing module, intelligently parse and unify various heterogeneous data formats and description methods, solve the structural and expression differences between systems, support cross-platform data integration, and thus greatly improve the verification efficiency and consistency. Relying on the semantic understanding and zero-shot text classification ability of large language models, the system can perform consistency checks on PD categories and descriptions within PD records, automatically discover and correct classification mismatches (such as "visit delay" being mislabeled as "medication error"), and adapt to different PD description scenarios without additional training data. Using the deep reasoning ability of large language models, the system can perform differential comparison verification on PD records in CTMS and EDC systems, comprehensively analyze structured fields and unstructured descriptions, identify subtle differences between records (such as semantic inconsistencies, misclassifications, or data missing), and infer potential root causes, providing accurate verification results and improvement suggestions, thereby improving data consistency and realizing the intelligence of the verification process.

[0170] Traditional methods (such as manual review or semi-automated processing) usually require a large amount of manual intervention, which is both time-consuming and inefficient. For example, in a large multi-center clinical trial involving thousands of PD records, manually comparing the data in CTMS and EDC often takes several days, while the method of this embodiment shortens the required time to a few minutes through automated verification, and the total time after manual confirmation can also be shortened to within 1 hour.

[0171] When dealing with unstructured PD descriptions using traditional methods, it often relies on manual judgment or simple rules, making it difficult to accurately identify records with similar expressions but different semantics, which in turn leads to inaccurate verification results. The method of this embodiment can verify data more accurately through intelligent means.

[0172] The method of this embodiment makes full use of the wide application of the Microsoft Office environment and adopts an operation interface familiar to users, thereby reducing the invasiveness of the system and ensuring seamless integration into the existing workflow. This design greatly reduces the learning cost of users and improves the compatibility of the system and the convenience of deployment.

[0173] The above-mentioned automatic verification method for clinical trial protocol deviations can efficiently and automatically process the consistency check and difference comparison of data by obtaining PD records and their classification and grading data from the CTMS system and the EDC system and using a large language model to parse the data; this process significantly improves the efficiency and accuracy of verification, ensures compatibility between different systems, and reduces manual intervention and learning costs. By generating a report containing consistency verification and difference comparison, the system can be seamlessly integrated into the existing workflow, realizing the automation and optimization of the process, and greatly improving work efficiency and quality.

[0174] Figure 9 It is a schematic block diagram of an automatic verification device 300 for clinical trial protocol deviations provided by an embodiment of the present invention. As Figure 9 shown, corresponding to the above automatic verification method for clinical trial protocol deviations, the present invention also provides an automatic verification device 300 for clinical trial protocol deviations. The automatic verification device 300 for clinical trial protocol deviations includes units for executing the above automatic verification method for clinical trial protocol deviations, and the device can be configured in a server. Specifically, please refer to Figure 9 , the automatic verification device 300 for clinical trial protocol deviations includes an acquisition unit 301, an analysis unit 302, a PD inspection unit 303, a report generation unit 304, and an output unit 305.

[0175] The acquisition unit 301 is used to acquire PD records and PD classification and grading data from the CTMS system and the EDC system to obtain two sets of data of the same project from the CTMS system and the EDC system; the analysis unit 302 is used to analyze the two sets of data using a large language model to obtain two analysis results; the PD inspection unit 303 is used to perform a consistency check on the two sets of data according to the two analysis results and compare the differences to obtain a detection result; the report generation unit 304 is used to generate a report containing consistency verification and difference comparison according to the detection result; the output unit 305 is used to output the report.

[0176] In one embodiment, the parsing unit 302 includes: A classification subunit, configured to classify the two pieces of data using a large language model to obtain deviation record data and classification and grading data; an information extraction subunit, configured to automatically map different system fields to the deviation record data and the classification and grading data and parse the unstructured PD description to extract key information; a conversion subunit, configured to convert the key information into a unified internal data structure to obtain two parsing results.

[0177] In one embodiment, the classification subunit is configured to split the two pieces of data into data blocks and classify each data block using the zero-shot classification ability of the large language model, and combine a voting mechanism to determine the final category to obtain deviation record data and classification and grading data.

[0178] In one embodiment, the information extraction subunit includes: A first construction module, configured to construct a prompt; a parsing module, configured to input the deviation record data and the classification and grading data into the large language model in combination with the prompt to parse the unstructured PD description to obtain a parsing result in JSON format; an unserialization module, configured to unserialize the parsing result in JSON format into an object of a programming language to obtain key information.

[0179] In one embodiment, the PD checking unit 303 includes: An internal consistency checking subunit, configured to check the internal consistency of the two parsing results through a large language model to obtain a consistency result; a comparison subunit, configured to compare the two parsing results by subject and time, mark missing records and evaluate the consistency of the two pieces of deviation record data to determine the missing mark and the reason for the difference; wherein, the detection result includes the consistency result, the missing mark and the reason for the difference.

[0180] In one embodiment, the internal consistency checking subunit includes: A second construction module, configured to construct a prompt according to the two parsing results; an analysis module, configured to input the two parsing results and the prompt into the large language model for internal consistency analysis to obtain a consistency result; wherein, the consistency result includes the consistency result, the confidence level, the reason for inconsistency and the suggestion.

[0181] In one embodiment, the comparison subunit includes: A grouping module for grouping the two parsing results according to the subject screening number, and classifying the records of the same subject according to the PD type to obtain multiple groups of data; a sorting module for sorting each group of data according to the event occurrence time to obtain a sorting result; a third construction module for constructing a comparison prompt word; a comparison module for inputting the sorting result and the comparison prompt word into a large language model to obtain a comparison result, where the comparison result includes the marker missing situation and the reason for the difference.

[0182] In one embodiment, the report generation unit 304 is used to generate a structured Excel file including consistency verification and difference comparison based on the detection result to obtain a report.

[0183] It should be noted that those skilled in the art can clearly understand the specific implementation processes of the above clinical trial protocol deviation automatic verification device 300 and each unit, which can refer to the corresponding descriptions in the foregoing method embodiments. For the convenience and conciseness of description, they will not be elaborated here.

[0184] The above clinical trial protocol deviation automatic verification device 300 can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 9 shown.

[0185] Please refer to Figure 9 , Figure 9 is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a server. Among them, the server can be an independent server or a server cluster composed of multiple servers.

[0186] Refer to Figure 9 , the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory can include a non-volatile storage medium 503 and an internal memory 504.

[0187] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, and when the program instructions are executed, the processor 502 can be caused to execute a clinical trial protocol deviation automatic verification method.

[0188] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0189] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be made to execute an automatic verification method for clinical trial protocol deviations.

[0190] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0191] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the following steps: Obtain PD records and PD classification and grading data from the CTMS system and the EDC system to obtain two sets of data for the same project in the CTMS system and the EDC system; use a large language model to parse the two sets of data to obtain two parsing results; perform consistency checks on the two sets of data according to the two parsing results, and compare the differences to obtain a detection result; generate a report containing consistency verification and difference comparison according to the detection result; output the report.

[0192] In one embodiment, when the processor 502 implements the step of using a large language model to parse the two sets of data to obtain two parsing results, the following steps are specifically implemented: Use a large language model to classify the two sets of data to obtain deviation record data and classification and grading data; automatically map different system fields to the deviation record data and classification and grading data and parse the unstructured PD description to extract key information; convert the key information into a unified internal data structure to obtain two parsing results.

[0193] In one embodiment, when the processor 502 implements the step of using a large language model to classify the two sets of data to obtain deviation record data and classification and grading data, the following steps are specifically implemented: Split the two sets of data into data blocks and use the zero-shot classification ability of the large language model to classify each data block, and combine the voting mechanism to determine the final category to obtain deviation record data and classification and grading data.

[0194] In one embodiment, when the processor 502 implements the step of automatically mapping different system fields to the deviation record data and classification and grading data and parsing the unstructured PD description to extract key information, the following steps are specifically implemented: Construct a prompt; input the deviation record data and classification and grading data combined with the prompt into a large language model to parse the unstructured PD description and obtain a parsed result in JSON format; deserialize the JSON format parsed result into an object of a programming language to obtain key information.

[0195] In one embodiment, when the processor 502 implements the step of performing a consistency check on the two parsed results and comparing the differences to obtain a detection result, the specific implementation is as follows: Check the internal consistency of the two parsed results through a large language model to obtain a consistency result; compare the two parsed results by subject and time, mark missing records, and evaluate the consistency of the two deviation record data to determine the missing marks and reasons for differences; wherein, the detection result includes a consistency result, missing marks, and reasons for differences.

[0196] In one embodiment, when the processor 502 implements the step of checking the internal consistency of the two parsed results through a large language model to obtain a consistency result, the specific implementation is as follows: Construct a prompt according to the two parsed results; input the two parsed results and the prompt into the large language model for internal consistency analysis to obtain a consistency result; Wherein, the consistency result includes a consistency result, confidence level, reasons for inconsistency, and suggestions.

[0197] In one embodiment, when the processor 502 implements the step of comparing the two parsed results by subject and time, marking missing records, and evaluating the consistency of the two deviation record data to determine the missing marks and reasons for differences, the specific implementation is as follows: Group the two parsed results according to the subject screening number, and classify the records under the same subject according to the PD type to obtain multiple groups of data; sort each group of data according to the event occurrence time to obtain a sorting result; construct a comparison prompt; input the sorting result and the comparison prompt into the large language model to obtain a comparison result, wherein the comparison result includes the marked missing situation and reasons for differences.

[0198] In one embodiment, when the processor 502 implements the step of generating a report including consistency verification and difference comparison based on the detection result, the specific implementation is as follows: Generate a structured Excel file including consistency verification and difference comparison based on the detection result to obtain a report.

[0199] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0200] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the flow steps of the embodiments of the above methods.

[0201] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor executes the following steps: Obtain PD records and PD classification and grading data from the CTMS system and the EDC system to obtain two sets of data for the same project in the CTMS system and the EDC system; use a large language model to parse the two sets of data to obtain two parsing results; perform consistency checks on the two sets of data according to the two parsing results and compare the differences to obtain a detection result; generate a report including consistency verification and difference comparison according to the detection result; output the report.

[0202] In one embodiment, when the processor executes the computer program to implement the step of using a large language model to parse the two sets of data to obtain two parsing results, the following steps are specifically implemented: Use a large language model to classify the two sets of data to obtain deviation record data and classification and grading data; automatically map different system fields for the deviation record data and the classification and grading data and parse the unstructured PD description to extract key information; convert the key information into a unified internal data structure to obtain two parsing results.

[0203] In one embodiment, when the processor executes the computer program to implement the step of classifying the two pieces of data using a large language model to obtain deviation record data and classification and grading data, the specific implementation is as follows: Split the two pieces of data into data blocks and classify each data block using the zero-shot classification ability of the large language model, and determine the final category in combination with a voting mechanism to obtain deviation record data and classification and grading data.

[0204] In one embodiment, when the processor executes the computer program to implement the step of automatically mapping different system fields to the deviation record data and classification and grading data and parsing the unstructured PD description to extract key information, the specific implementation is as follows: Construct a prompt; input the deviation record data and classification and grading data combined with the prompt into the large language model to parse the unstructured PD description to obtain a parsing result in JSON format; deserialize the JSON format parsing result into an object of a programming language to obtain key information.

[0205] In one embodiment, when the processor executes the computer program to implement the step of performing a consistency check on the two pieces of data according to the two parsing results, comparing the differences, and obtaining a detection result, the specific implementation is as follows: Check the internal consistency of the two parsing results through the large language model to obtain a consistency result; compare the two parsing results by subject and time, mark the missing records, and evaluate the consistency of the two pieces of deviation record data to determine the missing marks and reasons for the differences; Among them, the detection result includes a consistency result, missing marks, and reasons for the differences.

[0206] In one embodiment, when the processor executes the computer program to implement the step of checking the internal consistency of the two parsing results through the large language model to obtain a consistency result, the specific implementation is as follows: Construct a prompt according to the two parsing results; input the two parsing results and the prompt into the large language model for internal consistency analysis to obtain a consistency result; among them, the consistency result includes a consistency result, confidence level, reasons for inconsistency, and suggestions.

[0207] In one embodiment, when the processor executes the computer program to implement the step of comparing the two parsing results by subject and time, marking the missing records, and evaluating the consistency of the two pieces of deviation record data to determine the missing marks and reasons for the differences, the specific implementation is as follows: Group the two parsing results according to the subject screening number, and classify the records of the same subject according to the PD type to obtain multiple groups of data; sort each group of data according to the event occurrence time to obtain a sorting result; construct a comparison prompt word; input the sorting result and the comparison prompt word into a large language model to obtain a comparison result, where the comparison result includes the marker missing situation and the reason for the difference.

[0208] In one embodiment, when the processor executes the computer program to implement the step of generating a report including consistency verification and difference comparison according to the detection result, the following steps are specifically implemented: Generate a structured Excel file including consistency verification and difference comparison according to the detection result to obtain a report.

[0209] The storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.

[0210] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0211] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0212] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0213] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0214] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An automatic verification method for deviations in clinical trial protocols, characterized in that, It includes: Obtain PD records and PD classification and grading data from the CTMS system and the EDC system to get two sets of data for the same project in the CTMS system and the EDC system; Use a large language model to parse the two sets of data to obtain two parsing results; Perform consistency checks on the two sets of data based on the two parsing results, and compare the differences to obtain a detection result; Generate a report containing consistency verification and difference comparison based on the detection result; Output the report.

2. The automatic verification method for clinical trial protocol deviations according to claim 1, wherein The using a large language model to parse the two sets of data to obtain two parsing results includes: Use a large language model to classify the two sets of data to obtain deviation record data and classification and grading data; Automatically map different system fields to the deviation record data and the classification and grading data and parse the unstructured PD description to extract key information; Convert the key information into a unified internal data structure to obtain two parsing results.

3. The automatic verification method for clinical trial protocol deviations according to claim 2, characterized in that, The using a large language model to classify the two sets of data to obtain deviation record data and classification and grading data includes: Split the two sets of data into data blocks and use the zero-shot classification ability of the large language model to classify each data block, and combine the voting mechanism to determine the final category to obtain deviation record data and classification and grading data.

4. The automatic verification method for deviations from a clinical trial protocol according to claim 2, wherein The automatically mapping different system fields to the deviation record data and the classification and grading data and parsing the unstructured PD description to extract key information includes: Construct prompt words; Input the deviation record data and the classification and grading data combined with the prompt words into the large language model to parse the unstructured PD description to obtain a parsing result in JSON format; Deserialize the JSON format parsing result into an object in a programming language to obtain key information.

5. The automatic verification method for deviations from a clinical trial protocol according to claim 1, wherein The performing consistency checks on the two sets of data based on the two parsing results, and comparing the differences to obtain a detection result includes: Check the internal consistency of the two parsing results through the large language model to obtain a consistency result; Compare the two parsing results by subject and time, mark the missing records and evaluate the consistency of the two sets of deviation record data to determine the missing marks and reasons for differences; Among them, the detection result includes a consistency result, missing marks and reasons for differences.

6. The automatic verification method for clinical trial protocol deviations according to claim 5, wherein The checking the internal consistency of the two parsing results through the large language model to obtain a consistency result includes: Construct prompt words according to the two parsing results; Input the two parsing results and the prompt words into the large language model for internal consistency analysis to obtain a consistency result; Among them, the consistency result includes a consistency result, confidence level, reasons for inconsistency and suggestions.

7. The automatic verification method for deviations from a clinical trial protocol according to claim 5, characterized in that, The comparing the two parsing results by subject and time, marking the missing records and evaluating the consistency of the two sets of deviation record data to determine the missing marks and reasons for differences includes: Group the two parsing results according to the subject screening number, and classify the records under the same subject according to the PD type to obtain multiple groups of data; Sort each group of data according to the event occurrence time to obtain a sorting result; Construct a comparison prompt; Input the sorting result and the comparison prompt into a large language model to obtain a comparison result, where the comparison result includes the missing marker situation and the reason for the difference.

8. The automatic verification method for deviations from a clinical trial protocol according to claim 1, wherein Generating a report including consistency verification and difference comparison according to the detection result, including: Generating a structured Excel file including consistency verification and difference comparison according to the detection result to obtain a report.

9. An automatic verification device for deviations in clinical trial protocols, characterized in that, Including: An acquisition unit for acquiring PD records and PD classification and grading data from the CTMS system and the EDC system to obtain two sets of data of the same project in the CTMS system and the EDC system; A parsing unit for parsing the two sets of data using a large language model to obtain two parsing results; A PD inspection unit for performing consistency inspection on the two sets of data according to the two parsing results and comparing the differences to obtain a detection result; A report generation unit for generating a report including consistency verification and difference comparison according to the detection result; An output unit for outputting the report.

10. A computer device, characterized in that, The computer device includes a memory and a processor, and a computer program is stored on the memory. When the processor executes the computer program, the method described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Private data classification and grading method

    CN118133221A

  • Structured data generation model based on natural language processing

    CN118940719A

  • Method and system for carrying out standard labeling on case report form

    CN119517273A

  • Information extraction method for clinical test scheme

    CN119541882A

  • Medical data document classification and marking system

    CN119621972A

Cited By

  • Automatic data comparison method and device, terminal equipment and storage medium

    CN121456012A