Clinical test document quality detection and processing method and system, and computer equipment
Through the combination of file parser, OCR engine, BERT model and large language model, the automated processing of clinical trial documents is realized, solving the problems of low efficiency, poor consistency and high error rate in the existing technology, and improving the efficiency and accuracy of document processing.
Patent Information
- Application Number
- CN202510746148.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the quality detection and processing of clinical trial documents relies on artificial quality inspection, resulting in low efficiency, poor consistency and high error rate, making it difficult to meet the needs of modern clinical trials.
The file parser and OCR engine are used for format conversion and optimization processing, combined with the BERT model and the large language model for automatic classification, extract key information, and conduct integrity and compliance checks through the deep learning model, and finally realize automated storage and archiving.
It improves the efficiency and accuracy of document processing, reduces manual intervention, ensures the consistency and accuracy of documents, solves the problems of low efficiency and high error rate in the existing technology, and improves the overall processing quality.
Smart Images

Figure CN120257941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a document processing method, and more specifically to a method, system, and computer device for quality inspection and processing of clinical trial documents. Background Art
[0002] eTMF (electronic Trial Master File) is a digital system used to store, manage, and archive all documents required in clinical trials, which is an electronic upgrade of the traditional paper trial master file. eTMF covers various types of documents including clinical trial protocols, informed consent forms, investigator manuals, ethics committee approval documents, meeting minutes, data management plans, etc. Its core objective is to ensure the integrity, traceability, and compliance of clinical trial documents, so as to prove the legality of the trial and the reliability of the data during inspections by regulatory agencies.
[0003] The functions of the eTMF system include document upload and storage, metadata management, version control, audit trail, and compliance report generation, etc. The documents in the system are jointly maintained by clinical trial assistants, quality assurance personnel, research coordinators, etc. According to the requirements of ICH-GCP and FDA, the documents in the TMF must be kept "complete, accurate, timely, and readable". If there are quality problems with the documents, it may lead to compliance failures, which may in turn trigger legal proceedings or fines. If during the inspection by regulatory agencies, missing, unclear, or logically incorrect documents are found, the trial may be suspended, and even affect drug approval. As direct evidence of trial design, implementation, and results, if there are quality problems with clinical trial documents, it will lead to unreliable data, increase time costs, and may result in incorrect data analysis. Therefore, strict quality inspection and standardization processing must be carried out on the documents uploaded to the eTMF system.
[0004] Current document processing mainly relies on CTA (Clinical Trial Assistant) for manual quality inspection and processing. This manual process generally includes the following steps: The CTA receives scanned or paper documents and preliminarily confirms the quantity and type of the documents; Manually judge the type of the documents such as clinical trial protocols, subject informed consent forms, etc. according to the appearance or content of the documents; Manually read the documents, extract key information, and name the documents according to preset rules; Check the documents page by page to identify problems such as blank pages, unsigned pages, skewed pages, missing pages, etc.; Upload the processed documents to the eTMF system or folder.
[0005] However, manual quality inspection relies on subjective judgment. Different CTAs may have different understandings of quality standards, resulting in inconsistent results. It may take 1-2 minutes to process each document, and since the number of clinical trial documents is huge, a large amount of man-hours are consumed. Due to the large number of quality inspection items for each document, negligence may lead to missed inspections or misjudgments, with a relatively high error rate, possibly exceeding 50%. Moreover, for the processing method that extracts text using a single technical tool such as OCR, it is difficult to meet the requirements of modern clinical trials in terms of efficiency, consistency, and comprehensiveness.
[0006] Therefore, it is necessary to design a new method to achieve efficient and accurate document quality detection and automated processing, so as to improve the quality and efficiency of document processing and solve the problems of low efficiency, poor consistency, and high error rate in the existing technology. Summary of the Invention
[0007] The purpose of the present invention is to overcome the defects of the existing technology and provide a method, system, and computer device for quality detection and processing of clinical trial documents.
[0008] To achieve the above purpose, the present invention adopts the following technical solutions: A method for quality detection and processing of clinical trial documents, including: Obtain the clinical trial document to be processed; Perform format conversion and optimization processing on the clinical trial document to be processed to obtain a preprocessing result; Automatically classify the preprocessing result to obtain a classification result; Extract the key information of the document according to the classification result and generate a file name to obtain an extraction result; Check the integrity, compliance, and archivability of the corresponding document according to the extraction result, and perform automated processing according to the detection result to obtain an automated processing result; Store and archive the automated processing result.
[0009] A further technical solution thereof is: The performing format conversion and optimization processing on the clinical trial document to be processed to obtain a preprocessing result includes: Use a file parser to identify the document format of the clinical trial document to be processed and extract basic metadata to obtain an identification result; Split the clinical trial document to be processed according to the identification result to obtain a split result; Extract and optimize the scanned image of the split result to obtain an image extraction result; Use an OCR engine to extract text from the image extraction result and correct spelling and semantics to obtain a correction result; Convert the correction result into a standard format to obtain a preprocessing result.
[0010] Its further technical solution is as follows: Automatically classifying the preprocessing result to obtain a classification result, including: Using a pre-trained BERT model to perform a preliminary classification on the preprocessing result to obtain a preliminary classification result; When the preliminary classification result is a type that has not appeared or a type with a confidence level not meeting the requirements, using the zero-shot classification ability of the large language model to perform semantic inference on the preprocessing result to obtain a classification result; when the preliminary classification result is not a type that has not appeared and a type with a confidence level not meeting the requirements, then determining the preliminary classification result as the classification result.
[0011] Its further technical solution is as follows: When the preliminary classification result is a type that has not appeared or a type with a confidence level not meeting the requirements, using the zero-shot classification ability of the large language model to perform semantic inference on the preprocessing result to obtain a classification result, including: Inputting the preprocessing result into the large language model and using a preset prompt template for semantic reasoning, so that the large language model generates a classification label according to the context of the document to obtain a classification result.
[0012] Its further technical solution is as follows: Extracting the key information of the document according to the classification result and generating a file name to obtain an extraction result, including: Automatically extracting key information from the preprocessing result using the large language model according to the classification result; Generating a file name and metadata using a predetermined template according to the key information to obtain an extraction result.
[0013] Its further technical solution is as follows: Checking the integrity, compliance, and archivability of the corresponding document according to the extraction result and performing automated processing according to the detection result to obtain an automated processing result, including: Performing blank page detection, page orientation correction, empty item and unchecked item verification, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blurring detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain an automated processing result.
[0014] Its further technical solution is as follows: Performing blank page detection, page orientation correction, empty item and unchecked item verification, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blurring detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain an automated processing result, including: Using deep learning models, large language models, OCR, convolutional neural networks, and rule engine technologies to perform blank page detection, page orientation correction, verification of unchecked items, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color stamp judgment, page blur detection, scanned watermark recognition, and page integrity check based on the extraction results to obtain an automated processing result.
[0015] Its further technical solution is: The use of deep learning models, large language models, OCR, convolutional neural networks, and rule engine technologies to perform blank page detection, page orientation correction, verification of unchecked items, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color stamp judgment, page blur detection, scanned watermark recognition, and page integrity check based on the extraction results to obtain an automated processing result, including: Identify and mark pages with no content or blank in the extraction results; Analyze the geometric relationship of the text block coordinates in the extraction results to automatically identify and correct page tilt; Check whether required items in the document are missed in the extraction results through the rule engine; Extract page numbers from the extraction results and check the integrity and continuity of the page number sequence; Automatically adjust the document page order according to the extracted page numbers or metadata; Use a deep learning model to detect whether the document lacks a signature or seal in the extraction results; Use a CNN model to detect whether the stamp is in color in the extraction results. If it is not in color, prompt to rescan; Use a CNN model to analyze the image clarity in the extraction results and mark blurred pages; Use deep learning technology to detect watermark information in the document in the extraction results; Check the integrity of the header, footer, and page numbers in the extraction results.
[0016] The present invention also provides a clinical trial document quality detection and processing system, including: A document acquisition unit for acquiring clinical trial documents to be processed; A preprocessing unit for performing format conversion and optimization processing on the clinical trial documents to be processed to obtain a preprocessing result; A classification unit for automatically classifying the preprocessing results to obtain a classification result; An extraction unit for extracting key information of the document according to the classification result and generating a file name to obtain an extraction result; A quality inspection unit for checking the integrity, compliance, and archivability of corresponding documents according to the extraction results, and performing automated processing based on the inspection results to obtain an automated processing result; A storage and archiving unit for storing and archiving the automated processing result.
[0017] The present invention also provides a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the above-mentioned method is implemented.
[0018] The beneficial effects of the present invention compared with the prior art are as follows: By integrating advanced preprocessing, automatic classification, information extraction, quality inspection, and archiving technologies, the present invention realizes efficient and accurate processing of clinical trial documents. First, the system performs format conversion and optimization on the documents to be processed to improve the input quality; then, a deep learning model is used to automatically classify the document types, and combined with a large language model to extract key information and automatically generate a standardized file name; next, the quality inspection module checks the integrity, compliance, and archivability of the documents, automatically repairs problems and ensures that the documents meet the requirements; finally, all documents are accurately stored and archived after automated processing; this process greatly improves the efficiency of document processing, reduces manual intervention, ensures the consistency and accuracy of the documents, solves the problems of low efficiency and high error rate in the prior art, and improves the overall processing quality.
[0019] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic diagram of the application scenario of the clinical trial document quality inspection and processing method provided by the embodiment of the present invention; Figure 2 It is a schematic diagram of the process of the clinical trial document quality inspection and processing method provided by the embodiment of the present invention; Figure 3 It is a schematic diagram of the sub-process of the clinical trial document quality inspection and processing method provided by the embodiment of the present invention; Figure 4 It is a schematic diagram of the sub-process of the clinical trial document quality inspection and processing method provided by the embodiment of the present invention; Figure 5Schematic diagram of a sub - process of the clinical trial document quality detection and processing method provided by an embodiment of the present invention; Figure 6 Schematic diagram of a sub - process of the clinical trial document quality detection and processing method provided by an embodiment of the present invention; Figure 7 Schematic block diagram of the clinical trial document quality detection and processing system provided by an embodiment of the present invention; Figure 8 Schematic block diagram of the computer device provided by an embodiment of the present invention. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0023] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0024] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0025] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0026] Please refer to Figure 1 and Figure 2 , Figure 1 Schematic diagram of the application scenario of the clinical trial document quality detection and processing method provided by an embodiment of the present invention. Figure 2Schematic flowchart of the clinical trial document quality detection and processing method provided by the embodiments of the present invention. This clinical trial document quality detection and processing method is applied to a server, which interacts with terminals for data. This method realizes efficient and accurate quality detection and processing of clinical trial documents through a series of automated steps. First, a file parser is used for format conversion and optimization, and combined with an OCR engine and a deep learning model for text correction and image optimization. Then, the pre-trained BERT model and the zero-shot classification ability of the large language model are used to intelligently classify the documents, and the key information is automatically extracted by the large language model to generate file names. Further, the system comprehensively checks the integrity, compliance, and archivability of the documents, including blank page detection, page sorting adjustment, signature and seal detection, etc., and uses OCR, convolutional neural network (CNN), and deep learning technologies to ensure processing accuracy and consistency. This process significantly improves the quality and efficiency of document processing, solves the problems of low efficiency and high error rate in traditional methods, and ensures the accuracy and efficiency of automated processing results. It improves the efficiency and accuracy of document processing, ensures the quality of files, reduces manual dependence, and provides more reliable technical support for clinical trials.
[0027] Figure 2 is a schematic flowchart of the clinical trial document quality detection and processing method provided by the embodiments of the present invention. As Figure 2 shown, the method includes the following steps S110 to S160.
[0028] S110. Obtain the clinical trial documents to be processed.
[0029] In this embodiment, the system obtains the clinical trial documents to be processed through an interface or manual upload. These documents may include various formats, such as scanned PDFs, images (JPEG, PNG), electronic PDFs, etc. These formats of documents usually contain important clinical trial data and often require further standardization and processing.
[0030] The sources of the documents to be processed include the original documents uploaded by researchers or obtained from document management systems (such as eTMF, SharePoint) through other channels.
[0031] The document types can include trial protocols, informed consent forms, case report forms, etc. Each document type has different structural and content requirements.
[0032] S120. Perform format conversion and optimization processing on the clinical trial documents to be processed to obtain a preprocessing result.
[0033] In this embodiment, the preprocessing result is a standardized document after format conversion, optimization processing, and information extraction, which is convenient for subsequent classification and analysis.
[0034] Process and optimize the document to be processed to ensure that the document can be processed subsequently in a standardized format. The result of the preprocessing will be a standardized data structure, including the document that has been split, optimized, and key information extracted.
[0035] In one embodiment, please refer to Figure 3 , the above step S120 may include steps S121 to S125.
[0036] S121. Use a file parser to identify the document format of the clinical trial document to be processed and extract basic metadata to obtain an identification result.
[0037] In this embodiment, a file parser (such as Tika) is used to identify the format of the input clinical trial document and extract basic metadata. The file parser can automatically identify the document type (such as PDF, JPEG, PNG, etc.) according to the file header information and file extension. During the parsing process, the system will also extract the basic metadata of the file (such as file size, creation time, author, etc.), and these metadata will support the subsequent processing steps.
[0038] The file parser supports multiple file formats (PDF, JPEG, PNG, etc.). It can automatically determine the document type based on the file extension or file content and extract relevant metadata. For example, for a PDF file, the file parser will extract information such as the number of pages, author, and creation date of the document; for an image format, the parser will identify the size and resolution of the image.
[0039] The identification result includes the document type, basic metadata, and specific information that needs further processing.
[0040] S122. Split the clinical trial document to be processed according to the identification result to obtain a split result.
[0041] In this embodiment, multi-page documents are split. For multi-page documents such as PDF files or scanned images, the system will split them into independent single-page documents according to the structure of each page for subsequent processing.
[0042] For PDF documents, the system uses a page splitting algorithm (such as based on the PDF page tree structure) to split each page to generate a separate document for each page. For images or scanned files, the system will use image processing techniques such as OCR (Optical Character Recognition) to identify the content of each page. The module will automatically determine the splitting points of the page according to the file content and layout and generate new files for each page.
[0043] The final split result is to convert a multi-page document into multiple single-page documents (each page document with a unique identifier, such as document ID_page number), which is convenient for subsequent classification, information extraction, and quality inspection.
[0044] S123. Extract and optimize the scanned image from the split result to obtain an image extraction result.
[0045] In this embodiment, the image extraction result refers to the image in the split result that has been extracted and optimized.
[0046] For the scanned image, the system needs to perform image optimization processing to ensure that the quality of each image page is high enough for subsequent OCR recognition and quality detection.
[0047] The system uses image processing techniques (such as denoising, contrast enhancement, and skew correction) to optimize the scanned image. For example, denoising techniques (such as Gaussian filtering) can remove noise in the image, contrast enhancement (such as histogram equalization) can make the text content clearer, and skew correction (such as Hough transform) can correct the rotation problem of the scanned document. These processes ensure that the scanned image has sufficient clarity for subsequent OCR recognition and quality assessment.
[0048] Through these techniques, the quality of the scanned image is optimized, and the text, tables, and graphics in the image can be accurately recognized, providing high-quality input for subsequent OCR extraction.
[0049] S124. Use an OCR engine to extract text from the image extraction result and correct spelling and semantics to obtain a correction result.
[0050] In this embodiment, OCR (Optical Character Recognition) technology is used to extract text from scanned images or multi-page documents. OCR technology can not only recognize printed text but also recognize handwritten text (under certain conditions) and convert it into structured data for subsequent analysis and processing.
[0051] The OCR engine (such as PaddleOCR) extracts text from each page of the scanned document, supporting multi-language recognition (such as English, Chinese, etc.). By dividing the page into blocks, OCR can accurately recognize each text area and assign coordinates to each text area. The recognized text will undergo semantic correction to eliminate possible spelling or grammar errors and finally output structured data (such as JSON format), including text content, coordinates, and confidence values.
[0052] The extracted text includes not only the text content in the document but also the position information and credibility of each paragraph. This information structure provides a solid foundation for subsequent document classification, information extraction, and quality detection.
[0053] S125. Convert the correction result into a standard format to obtain a preprocessing result.
[0054] In this embodiment, format standardization processing is performed on all documents to be processed (whether PDF files, image files, or text after OCR extraction), ensuring that they enter the subsequent processing flow in a consistent format.
[0055] The system converts all input documents into a unified standard format (such as PDF / A or a standardized image format), while retaining the metadata of the documents. All documents will be converted into a common structure (such as PDF, image sequences, and OCR text), facilitating subsequent document classification, information extraction, and quality detection.
[0056] This step ensures that regardless of the original format of the input documents, all documents ultimately enter the system in a unified standard format, ensuring the smooth progress of subsequent steps.
[0057] S121 to S125 in the above step S120 cover a series of operations from document format recognition, splitting, optimization processing to OCR text extraction and format standardization. These operations together ensure that clinical trial documents can be processed efficiently and accurately in the system, providing a reliable data basis for subsequent document classification, information extraction, quality detection, etc. This process greatly improves the efficiency and accuracy of document processing and provides strong technical support for clinical trials. Using parallel processing technology, splitting and optimization are completed in seconds.
[0058] S130. Automatically classify the preprocessing result to obtain a classification result.
[0059] In this embodiment, the classification result refers to analyzing the preprocessed document through an automated model to determine its document type, such as a protocol, informed consent form, or case report form, etc.
[0060] Specifically, according to the preprocessing result of the document, automatically identify and determine the type of the document, providing accurate document classification information for subsequent modules such as information extraction, standardized naming, and quality detection. The accuracy of the classification result directly affects the efficiency and accuracy of the subsequent process.
[0061] In one embodiment, please refer to Figure 4 , the above step S130 may include steps S131 to S132.
[0062] S131. Use a pre-trained BERT model to perform a preliminary classification on the preprocessing result to obtain a preliminary classification result.
[0063] In this embodiment, the system uses the BERT (Bidirectional Encoder Representations from Transformers) model for preliminary classification. BERT is a pre-trained model based on the Transformer architecture and has very strong natural language understanding capabilities. BERT can capture the context information in the text and adapt to the document type classification task in a specific domain through fine-tuning.
[0064] The goal of the preliminary classification is to analyze the preprocessed document content through a pre-trained model and classify the document into several known common types (such as protocols, informed consent forms, case report forms, etc.). This process determines the document type by analyzing the text content in the preprocessing results (such as the text extracted by OCR). The classification of the document type is based on the labeled document dataset to ensure the classification accuracy.
[0065] The BERT model is fine-tuned on a training set containing thousands of labeled documents, which cover common clinical trial document types. The BERT model extracts the feature representations of the documents and outputs the probability values of each category through a classification head (usually a fully connected layer). The model outputs the most likely document type as the preliminary classification result. Due to the efficient text processing ability of the BERT model, the classification accuracy of this model can reach more than 98%, and it only takes about 0.05 seconds to complete the classification when processing documents.
[0066] The preliminary classification result includes the category of the document (such as "protocol" or "informed consent form").
[0067] If the document belongs to a known type, the system will directly adopt the preliminary classification result and continue with subsequent processing.
[0068] S132. When the preliminary classification result is a type that has not appeared or a type with a confidence level that does not meet the requirements, use the zero-shot classification ability of the large language model to perform semantic inference on the preprocessing result to obtain a classification result; when the preliminary classification result is not a type that has not appeared and a type with a confidence level that does not meet the requirements, then determine the preliminary classification result as the classification result.
[0069] In this embodiment, when the preliminary classification result is a type that has not appeared or a type with a confidence level that does not meet the requirements, input the preprocessing result into the large language model and use a preset prompt template for semantic reasoning, so that the large language model generates a classification label according to the context of the document to obtain a classification result.
[0070] In step S132, if an unappeared type occurs in the preliminary classification result, or the confidence level of the preliminary classification result does not meet the requirements (e.g., lower than 90%), the system will use the zero-shot classification ability of the large language model for semantic inference. This means that even without being specifically trained for a particular type, the system can understand the semantics of the document through the large language model and automatically infer its category.
[0071] Zero-shot learning is a classification method that does not require training data or labeled data. It relies on the deep semantic understanding ability of the large language model and can perform reasoning based on context and prompts.
[0072] In the classification task of clinical trial documents, when encountering a document type that has never appeared before, the system allows the large language model to make inferences by inputting the content of the document and clear prompts (e.g., "Is this document a protocol, an informed consent form, or another type?"). Based on the model's understanding of the text content, the appropriate document category is finally output.
[0073] If the confidence level of the preliminary classification model is lower than the threshold (e.g., 0.9), or the document belongs to a type that has not been seen before (e.g., some new document formats or special document types), zero-shot classification can use large language models (such as Qwen2.5 or DeepSeek) for dynamic supplementation.
[0074] This model can infer the type of the document through semantic understanding and give a new classification result, avoiding manual intervention and ensuring that the system can adapt to new document types.
[0075] The system designs predefined prompt templates to guide the large language model for inference. For example, "Is this a clinical trial protocol?" or "Is this a case report form?" etc. Through the reasoning ability of the large language model, the system can quickly respond and assign an appropriate type to the document. The classification result provided by the large language model includes not only the category but also the confidence level of the inference. The system can judge whether further processing or manual intervention is required based on this confidence level.
[0076] Through zero-shot classification, the system can efficiently process documents of unknown types and assign the correct category to them based on semantic inference.
[0077] Finally, the classification result of the document will be recorded and passed to subsequent processing links, such as information extraction, quality inspection and other modules.
[0078] Step S130 realizes the efficient and accurate classification of clinical trial documents by combining the BERT model and the zero-shot classification ability of large language models. In most cases, the BERT model can quickly and accurately complete the initial classification. When encountering unseen types or low confidence, zero-shot classification can provide more flexible and intelligent classification results through the inference ability of large language models. This automated classification method greatly improves the adaptability and accuracy of the system, providing a solid foundation for the subsequent processing of clinical trial documents.
[0079] S140. Extract the key information of the document according to the classification result and generate a file name to obtain the extraction result.
[0080] In this embodiment, the extraction result refers to the key information automatically extracted from the document and the file name and metadata generated according to a predetermined template.
[0081] In the process of clinical trial document processing, the process of extracting key information and generating a file name based on the document classification result. This process plays an important role in the document automation processing system, ensuring the efficiency and consistency of the document in processing, archiving, retrieval and other links.
[0082] The main objective of step S140 is to extract key information from the document content according to the classification result of the document, and use this information to generate standardized file names and metadata to support subsequent quality inspection, storage and archiving, etc.
[0083] In one embodiment, please refer to Figure 5 , the above step S140 may include steps S141 to S142.
[0084] S141. Automatically extract key information from the preprocessing result using a large language model according to the classification result.
[0085] In this embodiment, the system first uses a large language model to automatically extract the key information in the document according to the document classification result and in combination with the corresponding document type. The large language model has powerful natural language understanding ability and can accurately identify and extract specific fields from the text. The specific process is as follows: Classify the input clinical trial document to determine the type of the document (such as protocol, informed consent form, case report form, etc.). Based on this classification result, the system can further determine which key information needs to be extracted. Different types of documents contain different key information. For example: Protocol documents may need to extract information such as trial number, researcher signature, trial date, etc.
[0086] The informed consent form may involve fields such as signature status, patient number, signing date, etc.
[0087] The case report form needs to extract patient ID, trial phase, clinical observation results, etc.
[0088] Large language models (such as the pre-trained Qwen2.5 or DeepSeek) can automatically identify and extract these key information by combining prompt engineering with context semantic reasoning. For example, the system may use a prompt template: "Extract the trial number, research date, and signature status from this protocol". The large language model extracts relevant data fields from the text extracted by OCR through semantic understanding and supports document processing in different languages and formats.
[0089] The extracted information is output in the form of structured data, such as JSON format. Each extracted field will contain the corresponding field value and extraction confidence. In this way, the system can accurately convert the key information in the document into a machine-readable format for subsequent processing and storage.
[0090] S142. Generate a file name and metadata using a predetermined template based on the key information to obtain the extraction result.
[0091] In this embodiment, the system generates a standardized file name based on the key information extracted in S141 according to a predetermined naming template and generates metadata for the document. This step ensures the naming standardization, traceability, and archiving standardization of the document. The specific process is as follows: Based on the extracted key information (such as trial number, document type, date, etc.), the system will automatically generate a file name according to a predetermined naming template (such as "{trial number}{document type}{date}.pdf"). For example, if the extracted trial number is "CTR2023-001", the document type is "Protocol", and the date is "20231015", then the file name will be named "CTR2023-001_Protocol_20231015.pdf". This standardized file naming rule reduces the confusion of manual naming and improves the efficiency of document management.
[0092] In addition to the file name, the system will also generate metadata containing the key information of the document. These metadata include various attributes of the document, such as trial number, document type, creation date, version number, signature status, etc. The generated metadata is stored in JSON format and contains detailed information for each field. Metadata is not only helpful for the storage and retrieval of documents, but also provides support in the document quality inspection process, such as detecting whether there is a missing signature or whether the page numbers are complete.
[0093] The generated file names and metadata will be automatically stored in the specified storage system (such as eTMF system, SharePoint, etc.) together with the document files. The standardized storage of metadata ensures the compliance and traceability of documents in subsequent management, facilitating subsequent document retrieval and auditing.
[0094] Through steps S140, S141, and S142, the system can achieve efficient key information extraction and standardized naming during the automated processing of clinical trial documents. The application of large language models not only improves the accuracy of information extraction but also ensures the standardization and consistency of document naming, thus providing strong support for subsequent quality inspection and storage archiving. This process significantly improves the automation level and processing efficiency of clinical trial document processing, reduces manual intervention, and improves the quality and consistency of processing results.
[0095] S150. Check the integrity, compliance, and archivability of the corresponding document according to the extraction result, and perform automated processing according to the detection result to obtain an automated processing result.
[0096] In this embodiment, the key objective of this step is to ensure that the document meets the standards, is accurate, and can be smoothly archived and stored.
[0097] Specifically, perform blank page detection, page orientation correction, verification of unchecked empty items, page number verification, page sorting adjustment, detection of unsigned or unsealed pages, judgment of non-color seals, page blurriness detection, recognition of scanned watermarks, and inspection of page incompleteness according to the extraction result to obtain an automated processing result.
[0098] Adopt deep learning models, large language models, OCR, convolutional neural networks, and rule engine technologies to perform blank page detection, page orientation correction, verification of unchecked empty items, page number verification, page sorting adjustment, detection of unsigned or unsealed pages, judgment of non-color seals, page blurriness detection, recognition of scanned watermarks, and inspection of page incompleteness according to the extraction result to obtain an automated processing result.
[0099] In one embodiment, please refer to Figure 6 , the above step S150 may include steps S151 to S1510.
[0100] S151. Identify and mark pages with no content or blanks in the extraction result.
[0101] In this embodiment, the system automatically checks whether there are blank pages in the document. If the content of a certain page of the document is empty (the number of characters extracted by OCR is 0 or extremely small), it is marked as a blank page and processed, usually by deleting the page.
[0102] Specifically, through OCR (Optical Character Recognition) technology, the presence and quantity of characters on the page are detected. If the text content recognized by OCR is less than the set threshold (for example: 3 characters), the page is considered a blank page.
[0103] S152. Analyze the geometric relationship of the text block coordinates in the extraction result, and automatically identify and correct page tilt.
[0104] In this embodiment, the system automatically checks whether there is a page tilt problem in the document page, and performs page orientation correction as needed, so that the page layout is correct and convenient for subsequent document processing.
[0105] Use OCR-based page analysis technology. By analyzing the deviation between the angle of the page text block and the horizontal axis, determine whether the page is tilted. If the tilt angle exceeds a predetermined threshold (such as ±5°), the system automatically corrects the page orientation.
[0106] S153. Check whether required items in the document are missed in the extraction result through a rule engine.
[0107] In this embodiment, this function checks whether there are empty items that need to be filled or checked in the document, such as whether the signature box or check box on the informed consent form is filled.
[0108] Utilize the large language model combined with the information extracted by OCR to identify the items that need to be checked or filled, and match and check through the rule engine to see if there are any omissions. If a missing fill or check option is found, the system will mark it as abnormal and remind the user.
[0109] S154. Extract page numbers from the extraction result and check the integrity and continuity of the page number sequence.
[0110] In this embodiment, the system verifies the continuity of the page numbers in the document, ensures that the page numbers of each page of the document are not missing, and checks whether the document page numbers are correct.
[0111] OCR technology can extract the page numbers of each page of the document and compare them with the expected page number sequence. The system automatically identifies missing pages based on the extracted page number information. If a page number interruption is found (for example, page numbers are repeated or missing), the system will automatically generate a report and perform corresponding processing.
[0112] S155. Automatically adjust the document page order according to the extracted page numbers or metadata.
[0113] In this embodiment, if the page sorting is incorrect (such as discontinuous page numbers, or the order of the document content is chaotic), the system will automatically adjust the page order to ensure that the document is arranged in the correct order.
[0114] The system uses the page numbers extracted by OCR and document metadata (such as dates) to determine the order of document pages. By analyzing the page number sequence, the system automatically adjusts the order of the pages and generates an updated document.
[0115] S156. Use a deep learning model to detect whether the document is missing a signature or seal based on the extraction result.
[0116] In this embodiment, the system automatically checks whether there are parts in the document that must be signed or sealed. If it finds parts without signatures or seals, the system will mark these missing information and remind the user to supplement them.
[0117] By combining a convolutional neural network (CNN) with OCR technology, the system automatically detects the positions of signatures and seals in the document. For places where seals or signatures are missing, the system will generate a warning prompt.
[0118] S157. Use the CNN model to detect whether the seal is in color based on the extraction result. If it is not in color, prompt to rescan.
[0119] In this embodiment, this function detects whether the seal is in color. If the seal is in black and white or grayscale, the system will mark it as abnormal and remind to rescan the color version of the document.
[0120] Use a convolutional neural network (CNN) to detect the seal area in the image and analyze the distribution of its RGB channels. If it is found that the gray scale ratio of the seal color exceeds a set threshold (for example, more than 80%), then the seal is considered to be non - color, and the system prompts to rescan.
[0121] S158. Use the CNN model to analyze the image clarity of the extraction result and mark blurred pages.
[0122] In this embodiment, the system automatically detects whether the document pages are clear. If it finds that the pages are blurred, which may affect the accuracy of OCR recognition, the system will mark these pages.
[0123] The system uses a convolutional neural network (CNN) for blurred image detection. By training an annotated image dataset (clear / blurred), the CNN model can determine whether each page of the image is clear. If the image is not clear, the system will prompt the user to check the scanning quality.
[0124] S159. Use deep learning technology to detect watermark information in the document based on the extraction result.
[0125] In this embodiment, the system checks whether there is a watermark in the document (such as text watermarks like "copy" or "draft", etc.). If there is a watermark, the system will prompt that the document may not meet the requirements for final archiving.
[0126] Use deep learning object detection techniques (such as Faster R-CNN) to automatically identify the watermark text in the image and mark the location of the watermark. The system can determine whether the document complies with the filing specifications. If a watermark exists, the user will be reminded to rescan.
[0127] S1510. Check the integrity of the header, footer, and page numbers for the extraction result.
[0128] In this embodiment, the system checks the integrity of the page to ensure that information such as the document header, footer, and page numbers is not missing, and the content of each page is complete and available.
[0129] Use OCR and image analysis techniques to identify whether the header, footer, and page content of the document are intact. If there are missing or incomplete pages, the system will automatically identify and mark them, generating a report for the user to view.
[0130] The final result of the automated processing is a set of documents that have undergone comprehensive quality inspection and automated processing. According to the aforementioned inspection results, the documents will undergo the following types of processing: Delete blank pages; Automatically rotate and correct the page orientation; Mark and supplement missing signatures or checkboxes; Automatically adjust the page order; Generate a warning for pages without signatures or seals; Generate an inspection prompt for blurred pages or pages with watermarks; Remind to rescan documents with non-color seals; Through the above steps, the finally output documents will meet the quality standards of clinical trial documents and can be successfully filed without manual intervention, greatly improving the processing efficiency and accuracy, while also reducing errors and rework. These processing steps ensure that each document can meet regulatory requirements and quality requirements, and ensure a smooth entry into the storage and filing stage.
[0131] S160. Store and file the automated processing results.
[0132] In this embodiment, the documents that have undergone classification, naming, and quality inspection processing are automatically uploaded and filed to the specified storage system. Ensure the standardization and efficiency of the document filing process, especially in the clinical trial environment of large-scale document processing. Support multiple external storage platforms (such as SharePoint, eTMF systems, etc.) and seamlessly integrate with these systems through a unified interface. Through this module, it can be ensured that the storage structure of the documents complies with regulatory requirements and meets the compliance standards for clinical trial data filing.
[0133] Automatically upload documents to a preset storage path according to the document type and metadata. These storage paths follow certain specifications, can generate a reasonable directory structure based on the document type and trial information, and ensure that files are stored in the specified locations.
[0134] Ensure that file names, metadata, and storage structures comply with regulatory standards (such as FDA 21 CFR Part 11). The file names and storage structures of documents are standardized to ensure the consistency of each file during the archiving process and avoid management chaos caused by non-standard naming or inconsistent paths.
[0135] Integrate with external storage platforms (such as SharePoint, eTMF systems, etc.) through standardized API interfaces. Whether it is a traditional file storage system or a modern electronic storage management system, the storage and archiving module can be compatible with it to ensure seamless data flow and docking between different platforms.
[0136] Support seamless docking with external storage systems (such as SharePoint, eTMF systems) through standardized API interfaces. These interfaces are based on the RESTful architecture and support functions such as uploading, querying, and metadata synchronization to ensure secure data transmission and easy operation. The OAuth 2.0 authentication mechanism is used to ensure the security of data transmission.
[0137] Apply preset path rules according to the metadata of the document (such as trial number, document type, creation date, etc.) to automatically generate the target storage path. For example, according to the preset rules, the storage path of the document may be " / Trial_{trial number} / {document type} / ", which can ensure that the document can be quickly retrieved and accessed after uploading. Before uploading, the module will first check whether the target path exists, and if not, it will automatically create a folder. To improve efficiency and stability, the upload operation supports batch processing to ensure that a large number of documents can be uploaded simultaneously.
[0138] The storage of documents follows strict compliance standards, and file names, metadata, and storage structures meet regulatory requirements. The file name generation rule (such as "{trial number}{document type}{date}.pdf") can ensure the consistency of file naming and at the same time ensure the uniqueness of file names. Metadata is stored using the Dublin Core standard, including information such as the creation time, version number, and compliance mark of the document. These data are embedded in the PDF file or synchronized to the database of the external storage system. This structured storage method ensures the convenience of file management, auditing, tracking, etc.
[0139] Based on the metadata of the document, the storage and archiving module can dynamically calculate the target storage path. For example, for a document with a trial number of "CTR2023-001", its storage path may be automatically mapped to " / Trial_CTR2023-001 / Protocol / ". This dynamic path mapping not only improves the flexibility of the storage process but also avoids the risk of path errors and enhances the reliability of the storage operation. If the target path does not exist, the system will automatically create a folder and trigger a retry mechanism when the upload fails to ensure the successful upload of the document.
[0140] During the document upload process, compliance verification is performed to ensure that the file naming, metadata, and storage structure comply with regulatory requirements (such as FDA 21 CFR Part 11). Compliance verification can identify duplicate file names or path errors, ensuring the uniqueness of the files and avoiding storage failures or file management chaos caused by path errors or naming conflicts. Through this compliance check, the system can ensure the accuracy and integrity of the data throughout the storage and archiving process.
[0141] The method of this embodiment ensures the consistency and standardization of all documents during the storage process through automated and standardized path mapping and naming rules, reducing the errors that may be caused by manual operations. This not only improves the efficiency of document management but also avoids management problems caused by non-standard storage. Through the automated archiving function, the storage process of documents no longer depends on manual operations, greatly improving work efficiency. For large-scale clinical trials, especially in cases where a large number of documents need to be processed, the storage and archiving module can quickly and accurately complete the storage of documents, significantly shortening the overall processing time. Through compliance verification and automatically generating standard-compliant file naming rules, it is ensured that all archived documents meet regulatory requirements. This is crucial for the compliance management of clinical trials. Especially in the face of a large number of documents and data, automated compliance detection can effectively prevent compliance issues caused by human negligence.
[0142] Automation completes the entire process of document upload and storage, reducing manual intervention and lowering labor costs. The system does not require manual path configuration, naming, and storage checks, significantly reducing the operation complexity and labor input, while also reducing errors that may be caused by humans.
[0143] In summary, the automated processing of the storage and archiving of results, through technical means such as standardized path mapping, naming rules, and compliance verification, not only ensures the standardization and efficiency of the document storage process but also greatly improves the efficiency and accuracy of document archiving in the scenario of large-scale document processing, reduces the need for manual operations, and ultimately realizes the automation and high efficiency of clinical trial document processing.
[0144] The method of this embodiment consists of multiple steps, aiming to achieve the automated processing of clinical trial documents and improve efficiency and quality control. First is the preprocessing, which is responsible for standardizing documents in different formats, supporting operations such as splitting PDF files and image files, OCR text extraction, image denoising, and enhancement. Through parallel processing technology, this module can complete the splitting and optimization of documents within seconds, providing high-quality input for subsequent document classification and information extraction.
[0145] Secondly is document classification. Based on the BERT model and a text classification model trained with a large number of labeled clinical trial documents, it can accurately identify document types. For unseen document types, the system utilizes the zero-shot classification ability of the large language model to automatically classify them through semantic understanding, ensuring accurate classification even for non-standard documents. This module significantly improves the accuracy and generalization ability of document classification.
[0146] Next is information extraction and standardized naming. The large language model is used to automatically extract key information from the text extracted by OCR and generate standardized names. Whether it is an informed consent form, a case report form, or a trial protocol, the system can extract necessary fields according to the document category, ensure the unity of file naming, and provide marks for quality inspection, such as unsigned information, to ensure the smooth and accurate processing process.
[0147] The role of quality inspection is to ensure that the physical and format quality of documents meets the requirements. Through deep learning image processing technology, the system can detect problems such as blank pages, incorrect page orientation, missing pages, unsigned, duplicate pages, and blurred images, and perform automatic repair. For example, blank pages can be automatically deleted, and incorrect page orientation can also be corrected to ensure that the document quality meets the standard. This module greatly reduces the need for manual intervention and improves the accuracy of quality control.
[0148] Storage and archiving is responsible for automatically uploading the classified, named, and quality-inspected documents to the specified storage system. Using dynamic path mapping technology, this module can generate the correct storage path according to the document type and metadata, avoiding path errors. And it supports batch processing, significantly reducing the time and workload of manual uploading.
[0149] The entire method is designed to improve the efficiency and accuracy of document processing. Compared with traditional manual processing, the system can shorten the processing time of each document to 5-15 seconds. The processing time of large-scale clinical trial documents can be reduced to one-tenth of the original, significantly improving work efficiency. At the same time, the system greatly reduces manual intervention through automated processes, lowers the operating threshold and cost, avoids inconsistencies and errors in manual processing, and reduces the rework rate. Specifically, automation replaces the traditional process that relies on CTA manual quality inspection, reducing more than 90% of manual intervention, thereby significantly reducing manpower investment. Users only need to upload the original scan, and the system automatically completes all processing steps, eliminating professional skills requirements and complex configurations, greatly reducing the operating threshold. Through the preprocessing and quality inspection modules, the system fixes problems early (such as deleting blank pages, adjusting page orientation, etc.), avoiding later rework due to quality problems (such as rescanning or adding signatures).
[0150] The method of this embodiment not only improves the efficiency of document processing, but also ensures the quality of documents. The BERT model is used for document classification, a large language model is used for information extraction and naming normalization, and a deep learning model is combined for comprehensive quality testing. Finally, the system designs a unified interface to ensure compatibility with external systems, supports automatic archiving of documents, and improves the scalability and stability of the entire system.
[0151] The above-mentioned clinical trial document quality detection and processing method realizes efficient and accurate clinical trial document processing by integrating advanced preprocessing, automatic classification, information extraction, quality detection and archiving technologies. First, the system converts and optimizes the format of the documents to be processed to improve the input quality; then, a deep learning model is used to automatically classify the document type, and combined with a large language model to extract key information and automatically generate a standardized file name; then, the quality detection module is used to check the integrity, compliance and archivability of the document, automatically fix problems and ensure that the document meets the requirements; finally, all documents are accurately stored and archived after automated processing. This process greatly improves the efficiency of document processing, reduces manual intervention, ensures the consistency and accuracy of documents, solves the problems of low efficiency and high error rate in existing technologies, and improves the overall processing quality.
[0152] Figure 7 is a schematic block diagram of a clinical trial document quality detection and processing system 300 provided in an embodiment of the present invention. Figure 7 As shown, corresponding to the above clinical trial document quality detection and processing method, the present invention also provides a clinical trial document quality detection and processing system 300. The clinical trial document quality detection and processing system 300 includes a unit for executing the above clinical trial document quality detection and processing method, and the system can be configured in a server.Figure 7 , the clinical trial document quality inspection and processing system 300 includes a document acquisition unit 301, a preprocessing unit 302, a classification unit 303, an extraction unit 304, a quality inspection unit 305, and a storage and archiving unit 306.
[0153] The document acquisition unit 301 is used to acquire clinical trial documents to be processed; the preprocessing unit 302 is used to perform format conversion and optimization processing on the clinical trial documents to be processed to obtain a preprocessing result; the classification unit 303 is used to automatically classify the preprocessing result to obtain a classification result; the extraction unit 304 is used to extract key information of the document according to the classification result and generate a file name to obtain an extraction result; the quality inspection unit 305 is used to check the integrity, compliance, and archivability of the corresponding document according to the extraction result, and perform automated processing according to the detection result to obtain an automated processing result; the storage and archiving unit 306 is used to store and archive the automated processing result.
[0154] In one embodiment, the preprocessing unit 302 includes: An identification subunit, which is used to identify the document format of the clinical trial document to be processed by using a file parser and extract basic metadata to obtain an identification result; a splitting subunit, which is used to split the clinical trial document to be processed according to the identification result to obtain a splitting result; an image extraction subunit, which is used to extract and optimize the scanned image of the splitting result to obtain an image extraction result; a text extraction subunit, which is used to extract text from the image extraction result by using an OCR engine and correct spelling and semantics to obtain a correction result; a conversion subunit, which is used to convert the correction result into a standard format to obtain a preprocessing result.
[0155] In one embodiment, the classification unit 303 includes: A preliminary classification subunit, which is used to perform preliminary classification on the preprocessing result by using a pre-trained BERT model to obtain a preliminary classification result; a category determination subunit, which is used to perform semantic inference on the preprocessing result by using the zero-shot classification ability of a large language model when the preliminary classification result is a type that has not appeared or the confidence level does not meet the requirements to obtain a classification result; when the preliminary classification result is not a type that has not appeared and the confidence level does not meet the requirements, the preliminary classification result is determined as the classification result.
[0156] In one embodiment, the category determination subunit is used to input the preprocessing result into a large language model and perform semantic reasoning by using a preset prompt word template, so that the large language model generates a classification label according to the context of the document to obtain a classification result.
[0157] In one embodiment, the extraction unit 304 includes: An information extraction subunit, configured to automatically extract key information from the preprocessing result using a large language model according to the classification result; and a generation subunit, configured to generate a file name and metadata using a predetermined template according to the key information to obtain an extraction result.
[0158] In one embodiment, the quality detection unit 305 is configured to perform blank page detection, page orientation correction, empty item and missed check verification, page number verification, page sorting adjustment, unsigned and unsealed detection, unsealed non-color judgment, page blurriness detection, scanned watermark recognition, and page incompleteness check on the extraction result to obtain an automated processing result.
[0159] In one embodiment, the quality detection unit 305 is configured to perform blank page detection, page orientation correction, empty item and missed check verification, page number verification, page sorting adjustment, unsigned and unsealed detection, unsealed non-color judgment, page blurriness detection, scanned watermark recognition, and page incompleteness check on the extraction result using a deep learning model, a large language model, OCR, a convolutional neural network, and rule engine technology to obtain an automated processing result.
[0160] In one embodiment, the quality detection unit 305 includes: A page recognition and marking subunit, configured to recognize and mark pages with no content or blank in the extraction result; an angle analysis subunit, configured to analyze the geometric relationship of the text block coordinates in the extraction result to automatically recognize and correct page tilt; a missed selection detection subunit, configured to check whether required items in the document are missed in the extraction result through a rule engine; a page number detection subunit, configured to extract page numbers from the extraction result and check the integrity and continuity of the page number sequence; a sequence adjustment subunit, configured to automatically adjust the document page sequence according to the extracted page numbers or metadata; a signature and seal detection subunit, configured to detect whether the document is missing a signature or seal in the extraction result using a deep learning model; a seal detection subunit, configured to detect whether the seal is in color in the extraction result through a CNN model, and prompt re-scanning if it is non-color; a page marking subunit, configured to analyze the image clarity of the extraction result using a CNN model and mark blurry pages; a watermark extraction subunit, configured to detect watermark information in the document in the extraction result using deep learning technology; and an integrity detection subunit, configured to check the integrity of the header, footer, and page numbers in the extraction result.
[0161] It should be noted that those skilled in the art can clearly understand the specific implementation processes of the above clinical trial document quality detection and processing system 300 and each unit, which can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity of description, they will not be elaborated here.
[0162] The above-mentioned clinical trial document quality detection and processing system 300 can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 8 .
[0163] Please refer to Figure 8 . Figure 8 FIG. is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 is a server. Among them, the server can be an independent server or a server cluster composed of multiple servers.
[0164] Refer to Figure 8 . The computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.
[0165] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions. When the program instructions are executed, the processor 502 can be caused to execute a method for detecting and processing the quality of clinical trial documents.
[0166] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0167] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be caused to execute a method for detecting and processing the quality of clinical trial documents.
[0168] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 8 the structure shown in FIG. is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0169] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the following steps: Obtain the clinical trial document to be processed; perform format conversion and optimization processing on the clinical trial document to be processed to obtain a preprocessing result; perform automatic classification on the preprocessing result to obtain a classification result; extract key information of the document according to the classification result and generate a file name to obtain an extraction result; check the integrity, compliance, and archivability of the corresponding document according to the extraction result, and perform automated processing according to the detected result to obtain an automated processing result; store and archive the automated processing result.
[0170] In one embodiment, when the processor 502 implements the step of performing format conversion and optimization processing on the clinical trial document to be processed to obtain a preprocessing result, the specific implementation is as follows: Use a file parser to identify the document format of the clinical trial document to be processed and extract basic metadata to obtain an identification result; split the clinical trial document to be processed according to the identification result to obtain a split result; extract and optimize the scanned image of the split result to obtain an image extraction result; use an OCR engine to extract text from the image extraction result and correct spelling and semantics to obtain a correction result; convert the correction result to a standard format to obtain a preprocessing result.
[0171] In one embodiment, when the processor 502 implements the step of performing automatic classification on the preprocessing result to obtain a classification result, the specific implementation is as follows: Use a pre-trained BERT model to perform preliminary classification on the preprocessing result to obtain a preliminary classification result; when the preliminary classification result is a type that does not appear or the confidence level does not meet the requirements, use the zero-shot classification ability of the large language model to perform semantic inference on the preprocessing result to obtain a classification result; when the preliminary classification result is not a type that does not appear and the confidence level does not meet the requirements, determine the preliminary classification result as the classification result.
[0172] In one embodiment, when the processor 502 implements the step of when the preliminary classification result is a type that does not appear or the confidence level does not meet the requirements, use the zero-shot classification ability of the large language model to perform semantic inference on the preprocessing result to obtain a classification result, the specific implementation is as follows: Input the preprocessing result into the large language model and perform semantic reasoning using a preset prompt template, so that the large language model generates a classification label according to the context of the document to obtain a classification result.
[0173] In one embodiment, when the processor 502 implements the step of extracting key information of the document according to the classification result and generating a file name to obtain an extraction result, the specific implementation is as follows: Automatically extract key information from the preprocessing result using a large language model according to the classification result; generate a file name and metadata using a predetermined template according to the key information to obtain an extraction result.
[0174] In one embodiment, when the processor 502 implements the step of checking the integrity, compliance, and archivability of the corresponding document according to the extraction result and performing automated processing according to the detected result to obtain an automated processing result, the specific implementation is as follows: Perform blank page detection, page orientation correction, empty item missing check, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blur detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain an automated processing result.
[0175] In one embodiment, when the processor 502 implements the step of performing blank page detection, page orientation correction, empty item missing check, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blur detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain an automated processing result, the specific implementation is as follows: Use deep learning models, large language models, OCR, convolutional neural networks, and rule engine technologies to perform blank page detection, page orientation correction, empty item missing check, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blur detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain an automated processing result.
[0176] In one embodiment, when the processor 502 implements the step of using deep learning models, large language models, OCR, convolutional neural networks, and rule engine technologies to perform blank page detection, page orientation correction, empty item missing check, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blur detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain an automated processing result, the specific implementation is as follows: Identify and mark the pages with no content or blanks in the extraction results; analyze the geometric relationships of the text block coordinates in the extraction results to automatically identify and correct page inclination; check whether required items in the document are missed in the extraction results through a rule engine; extract page numbers from the extraction results and check the integrity and continuity of the page number sequence; automatically adjust the document page order according to the extracted page numbers or metadata; use a deep learning model to detect whether the document in the extraction results lacks a signature or seal; use a CNN model to detect whether the seal in the extraction results is in color, and if it is not in color, prompt to rescan; use a CNN model to analyze the image clarity of the extraction results and mark blurred pages; use deep learning technology to detect watermark information in the document in the extraction results; check the integrity of the header, footer, and page numbers in the extraction results.
[0177] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0178] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the flow steps of the embodiments of the above methods.
[0179] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the following steps: Obtain the clinical trial document to be processed; perform format conversion and optimization processing on the clinical trial document to be processed to obtain a preprocessing result; perform automatic classification on the preprocessing result to obtain a classification result; extract key information from the document according to the classification result and generate a file name to obtain an extraction result; check the integrity, compliance, and archivability of the corresponding document according to the extraction result, and perform automated processing according to the detection result to obtain an automated processing result; store and archive the automated processing result.
[0180] In one embodiment, when the processor executes the computer program to implement the step of performing format conversion and optimization processing on the clinical trial document to be processed to obtain a preprocessing result, the following steps are specifically implemented: Use a file parser to identify the document format of the clinical trial document to be processed and extract basic metadata to obtain an identification result; split the clinical trial document to be processed according to the identification result to obtain a split result; extract and optimize the scanned image of the split result to obtain an image extraction result; use an OCR engine to extract text from the image extraction result and correct spelling and semantics to obtain a correction result; convert the correction result into a standard format to obtain a preprocessing result.
[0181] In one embodiment, when the processor executes the computer program to implement the step of performing automatic classification on the preprocessing result to obtain a classification result, the following steps are specifically implemented: Use a pre-trained BERT model to perform a preliminary classification on the preprocessing result to obtain a preliminary classification result; when the preliminary classification result is a type that has not appeared or the confidence level does not meet the requirements, use the zero-shot classification ability of the large language model to perform semantic inference on the preprocessing result to obtain a classification result; when the preliminary classification result is not a type that has not appeared and the confidence level does not meet the requirements, then determine the preliminary classification result as the classification result.
[0182] In one embodiment, when the processor executes the computer program to implement the step of when the preliminary classification result is a type that has not appeared or the confidence level does not meet the requirements, use the zero-shot classification ability of the large language model to perform semantic inference on the preprocessing result to obtain a classification result, the following steps are specifically implemented: Input the preprocessing result into the large language model and perform semantic reasoning using a preset prompt template, so that the large language model generates a classification label according to the context of the document to obtain a classification result.
[0183] In one embodiment, when the processor executes the computer program to implement the step of extracting key information of the document according to the classification result and generating a file name to obtain the extraction result, the specific implementation is as follows: Automatically extract key information from the preprocessing result using a large language model according to the classification result; generate a file name and metadata according to the key information using a predetermined template to obtain the extraction result.
[0184] In one embodiment, when the processor executes the computer program to implement the step of checking the integrity, compliance, and archivability of the corresponding document according to the extraction result and performing automated processing according to the detected result to obtain the automated processing result, the specific implementation is as follows: Perform blank page detection, page orientation correction, empty item missing check, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blur detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain the automated processing result.
[0185] In one embodiment, when the processor executes the computer program to implement the step of performing blank page detection, page orientation correction, empty item missing check, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blur detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain the automated processing result, the specific implementation is as follows: Use deep learning models, large language models, OCR, convolutional neural networks, and rule engine technologies to perform blank page detection, page orientation correction, empty item missing check, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blur detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain the automated processing result.
[0186] In one embodiment, when the processor executes the computer program to implement the step of using deep learning models, large language models, OCR, convolutional neural networks, and rule engine technologies to perform blank page detection, page orientation correction, empty item missing check, page number verification, page sorting adjustment, unsigned and unsealed detection, non-color seal judgment, page blur detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain the automated processing result, the specific implementation is as follows: Identify and mark pages with no content or blanks in the extraction results; analyze the geometric relationships of the text block coordinates in the extraction results to automatically identify and correct page tilt; check whether required items in the document are missed in the extraction results through a rule engine; extract page numbers from the extraction results and check the integrity and continuity of the page number sequence; automatically adjust the document page order according to the extracted page numbers or metadata; use a deep learning model to detect whether the document in the extraction results lacks a signature or seal; use a CNN model to detect whether the seal in the extraction results is in color, and if it is not in color, prompt to rescan; use a CNN model to analyze the image clarity in the extraction results and mark blurred pages; use deep learning technology to detect watermark information in the document in the extraction results; check the integrity of the header, footer, and page numbers in the extraction results.
[0187] The storage medium may be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk, an optical disk, or other various computer-readable storage media that can store program codes.
[0188] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0189] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0190] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the system embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0191] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0192] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for quality inspection and processing of clinical trial documents, characterized in that, Including: Obtain the clinical trial document to be processed; Perform format conversion and optimization processing on the clinical trial document to be processed to obtain a preprocessing result; Automatically classify the preprocessing result to obtain a classification result; Extract the key information of the document according to the classification result and generate a file name to obtain an extraction result; Check the integrity, compliance, and archivability of the corresponding document according to the extraction result, and perform automated processing according to the detected result to obtain an automated processing result; Store and archive the automated processing result.
2. The quality inspection and processing method of the clinical trial document according to claim 1, wherein The performing format conversion and optimization processing on the clinical trial document to be processed to obtain a preprocessing result includes: Use a file parser to identify the document format and extract basic metadata from the clinical trial document to be processed to obtain an identification result; Split the clinical trial document to be processed according to the identification result to obtain a split result; Extract and optimize the scanned image of the split result to obtain an image extraction result; Use an OCR engine to extract text from the image extraction result and correct spelling and semantics to obtain a correction result; Convert the correction result to a standard format to obtain a preprocessing result.
3. The quality inspection and processing method of clinical trial documents according to claim 1, characterized in that The automatically classifying the preprocessing result to obtain a classification result includes: Use a pre-trained BERT model to perform a preliminary classification on the preprocessing result to obtain a preliminary classification result; When the preliminary classification result is a type that has not appeared or a type with a confidence level that does not meet the requirements, use the zero-shot classification ability of the large language model to perform semantic inference on the preprocessing result to obtain a classification result; when the preliminary classification result is not a type that has not appeared and a type with a confidence level that does not meet the requirements, then determine the preliminary classification result as the classification result.
4. The quality inspection and processing method of clinical trial documents according to claim 3, characterized in that The when the preliminary classification result is a type that has not appeared or a type with a confidence level that does not meet the requirements, use the zero-shot classification ability of the large language model to perform semantic inference on the preprocessing result to obtain a classification result includes: Input the preprocessing result into the large language model and use a preset prompt template for semantic reasoning, so that the large language model generates a classification label according to the context of the document to obtain a classification result.
5. The quality inspection and processing method of the clinical trial document according to claim 1, characterized in that The extracting the key information of the document according to the classification result and generating a file name to obtain an extraction result includes: Automatically extract key information from the preprocessing result using a large language model according to the classification result; Generate a file name and metadata using a predetermined template according to the key information to obtain an extraction result.
6. The quality inspection and processing method for clinical trial documents according to claim 1, characterized in that The checking the integrity, compliance, and archivability of the corresponding document according to the extraction result, and performing automated processing according to the detected result to obtain an automated processing result includes: Perform blank page detection, page orientation correction, empty item missing check, page number verification, page sorting adjustment, unsigned seal detection, non-color seal judgment, page blur detection, scanned watermark recognition, and page incompleteness check according to the extraction result to obtain an automated processing result.
7. The quality inspection and processing method of clinical trial documents according to claim 6, characterized in that Performing blank page detection, page orientation correction, unchecked item verification, page number verification, page sorting adjustment, unsigned seal detection, non-color seal judgment, page blurriness detection, scanned watermark recognition, and page incompleteness check based on the extraction result to obtain an automated processing result, including: Using deep learning models, large language models, OCR, convolutional neural networks, and rule engine technologies to perform blank page detection, page orientation correction, unchecked item verification, page number verification, page sorting adjustment, unsigned seal detection, non-color seal judgment, page blurriness detection, scanned watermark recognition, and page incompleteness check based on the extraction result to obtain an automated processing result.
8. The quality inspection and processing method of the clinical trial document according to claim 7, characterized in that The using of deep learning models, large language models, OCR, convolutional neural networks, and rule engine technologies to perform blank page detection, page orientation correction, unchecked item verification, page number verification, page sorting adjustment, unsigned seal detection, non-color seal judgment, page blurriness detection, scanned watermark recognition, and page incompleteness check based on the extraction result to obtain an automated processing result, including: Identifying and marking pages with no content or blank in the extraction result; Analyzing the geometric relationship of the text block coordinates in the extraction result to automatically identify and correct page tilt; Checking whether required items in the document are missed in the extraction result through a rule engine; Extracting page numbers from the extraction result and checking the integrity and continuity of the page number sequence; Automatically adjusting the document page order according to the extracted page numbers or metadata; Using a deep learning model to detect whether the document is missing a signature or seal in the extraction result; Detecting whether the seal is in color through a CNN model for the extraction result, and if it is non-color, prompting to rescan; Using a CNN model to analyze the image clarity of the extraction result and marking blurry pages; Using deep learning technology to detect watermark information in the document based on the extraction result; Checking the integrity of the header, footer, and page numbers in the extraction result.
9. A clinical trial document quality inspection and processing system, characterized in that Including: A document acquisition unit for acquiring a clinical trial document to be processed; A preprocessing unit for performing format conversion and optimization processing on the clinical trial document to be processed to obtain a preprocessing result; A classification unit for automatically classifying the preprocessing result to obtain a classification result; An extraction unit for extracting key information of the document according to the classification result and generating a file name to obtain an extraction result; A quality detection unit for checking the integrity, compliance, and archivability of the corresponding document according to the extraction result and performing automated processing according to the detection result to obtain an automated processing result; A storage and archiving unit for storing and archiving the automated processing result.
10. A computer device, characterized in that, The computer device includes a memory and a processor, and a computer program is stored on the memory. When the processor executes the computer program, the method described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Archive digital duplicate quality automatic detection method
CN107194659A
Intelligent classification method and device for electronic files, electronic equipment and storage medium
CN110188077A
Method for enhancing document processing flow based on large language model
CN118551046A
Electronic file data quality detection method and system
CN119226280A
Medical document automatic generating and filing system and method
CN119479971A
Cited By
Resource consumption intelligent supervision system and method for document data quality detection
CN120562379A
Electronic document refined naming method and computer equipment
CN120849685A
Document auditing method and device based on cooperation of large model and rule engine
CN121435956A
Document auditing method and device based on cooperation of large model and rule engine
CN121435956B