Clinical test sensitive data automatic shielding method and system based on mobile terminal
By integrating camera cropping, transfer learning classification, OCR technology and fine-tuning BERT model on mobile devices, efficient and precise occlusion of sensitive data in clinical trials is achieved, inefficiency and compliance issues in the existing technology are solved, and data security and privacy are ensured.
Patent Information
- Application Number
- CN202510768491.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
AI Technical Summary
In clinical trials, the privacy protection methods for sensitive data on mobile devices are inefficient, inaccurate, and do not comply with the regulations that data must not leave the hospital.
The mobile device camera is used to obtain images, detect document borders and crop them through machine learning, combine the classification of convolutional neural network model of transfer learning, and extract text content using OCR technology. The fine-tuned BERT model is used to identify sensitive entities and perform irreversible occlusion operations.
It realizes efficient and accurate sensitive data obscuration on mobile devices, ensures data security and privacy, complies with privacy protection regulations, and reduces the risks of manual operations and server-side processing.
Smart Images

Figure CN120279560A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data processing method, and more particularly to an automatic masking method and system for clinical trial sensitive data based on a mobile device. Background Art
[0002] During the clinical trial process, to ensure the scientificity and safety of new drug research and development and medical technology verification, a large amount of sensitive data management is involved. This data includes personal information of patients and medical staff, such as key information like names, contact information, ID numbers, etc. When clinical research coordinators use mobile devices to digitally record various documents, they need to follow strict privacy protection standards to ensure complete desensitization and irreversibility during subsequent data processing.
[0003] Current privacy protection solutions mainly include three methods: manual smearing, rule - based automatic processing, and server - side processing. However, they each have problems such as low efficiency, insufficient accuracy, or compliance issues. For example, manual smearing is time - consuming and laborious and it is difficult to ensure complete masking of all sensitive information; rule - based methods can only identify information in a fixed format and cannot handle complex and diverse privacy content; while server - side processing does not meet the requirement of data not leaving the hospital because data needs to be transmitted to an external server.
[0004] Therefore, it is necessary to design a new method that can effectively mask sensitive data on a mobile device with limited resources while meeting the requirements of efficient and accurate processing, so as to adapt to strict privacy protection needs, ensure data security and privacy, and at the same time have high operational convenience and real - time performance. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the prior art and provide an automatic masking method and system for clinical trial sensitive data based on a mobile device.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: An automatic masking method for clinical trial sensitive data based on a mobile device, including: Obtaining an image containing clinical trial data through a mobile device camera, automatically detecting the document border and cropping out the effective area to obtain an initial image; Using a convolutional neural network model trained by transfer learning to classify the initial image to distinguish between a paper document and a computer screen screenshot, and selecting a corresponding processing strategy according to the classification result; Processing the initial image according to the processing strategy, and extracting the text content and position information in the processed image through OCR technology to obtain an extraction result; Inputting the extraction result into a fine - tuned BERT model to identify sensitive entities by combining context semantics to obtain the position information of the sensitive entities; Perform an irreversible masking operation on the corresponding region in the image containing clinical trial data according to the location information of the sensitive entity to obtain a desensitized image; Store or output the desensitized image.
[0007] A further technical solution thereof is that: obtaining an image containing clinical trial data through a mobile device camera, automatically detecting the document border and cropping out the valid region to obtain an initial image, including: Obtain an image containing clinical trial data through a mobile device camera; Apply machine learning technology to automatically identify the document edge of the image and generate a bounding box; Crop the image based on the determined bounding box coordinates, and use the OpenCV tool to extract and optimize the cropped region to obtain an initial image.
[0008] A further technical solution thereof is that: the training process of the convolutional neural network model includes: Collect various types of desensitized images from the historical records; Use LabelStudio to classify and label the collected images, and divide the labeled images to obtain a sample set; Select YOLOv11m as the base model, and train the base model through transfer learning combined with the sample set; Export the trained model in ONNX format.
[0009] A further technical solution thereof is that: processing the initial image according to the processing strategy, and extracting the text content and location information in the processed image through OCR technology to obtain an extraction result, including: Process the initial image belonging to a paper document image through denoising, contrast enhancement and binarization, remove moiré patterns in the initial image belonging to a computer screen image, and perform denoising, contrast enhancement and binarization to obtain a processing result; Perform horizontal line and vertical line recognition on the processing result to obtain a recognition result; Extract the text content and location information in the processing result through OCR technology according to the recognition result to obtain an extraction result.
[0010] A further technical solution thereof is that: extracting the text content and location information in the processing result through OCR technology according to the recognition result to obtain an extraction result, including: Identify and sort the text content and location information in the processing result using the PP-OCRv4 model on ONNX Runtime according to the recognition result to obtain an extraction result.
[0011] Its further technical solution is: The training process of the fine-tuned BERT model includes: Select a lightweight MobileBERT model; Construct and annotate a clinical trial text dataset; Enhance the clinical trial text dataset through synonym replacement, entity forgery, noise injection, and format variation strategies to obtain rich results; Use the rich results to fine-tune the MobileBERT model for specific sensitive entity recognition tasks based on sequence annotation tasks using the cross-entropy loss function and the AdamW optimizer to obtain the fine-tuned BERT model; Export the fine-tuned BERT model in ONNX format.
[0012] Its further technical solution is: Input the extraction result into the fine-tuned BERT model, and combine the context semantics to recognize sensitive entities to obtain the position information of the sensitive entities, including: After processing the extraction result through segmentation and encoding, input it into the fine-tuned BERT model, and combine the context semantics to recognize sensitive entities to obtain the position information of the sensitive entities.
[0013] Its further technical solution is: According to the position information of the sensitive entities, perform an irreversible masking operation on the corresponding area in the image containing clinical trial data to obtain a desensitized image, including: Use black filling and apply 3×3 Gaussian blur to the edge or perform manual masking. According to the position information of the sensitive entities, perform an irreversible masking operation on the corresponding area in the image containing clinical trial data to obtain a desensitized image.
[0014] The present invention also provides a mobile-based automatic masking system for clinical trial sensitive data, including: An image acquisition unit, used to acquire an image containing clinical trial data through a mobile device camera, automatically detect the document border and crop out the effective area to obtain an initial image; A classification unit, used to classify the initial image using a convolutional neural network model trained by transfer learning to distinguish between paper documents and computer screen screenshots, and select corresponding processing strategies according to the classification results; A processing unit, used to process the initial image according to the processing strategy, and extract the text content and position information in the processed image through OCR technology to obtain an extraction result; A recognition unit, used to input the extraction result into the fine-tuned BERT model, and combine the context semantics to recognize sensitive entities to obtain the position information of the sensitive entities; A masking unit for performing an irreversible masking operation on a corresponding region in an image containing clinical trial data according to the location information of the sensitive entity to obtain a desensitized image; A storage and output unit for storing or outputting the desensitized image.
[0015] The present invention also provides a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the above method is implemented.
[0016] The beneficial effects of the present invention compared with the prior art are as follows: By integrating a series of advanced technologies, the present invention realizes efficient and accurate masking of sensitive clinical trial data on mobile devices, while ensuring data security and privacy. First, the device camera is used to obtain an image and automatically detect and crop the document border to extract the effective region. Then, a convolutional neural network model trained by transfer learning is used to classify the image, and a suitable processing strategy is selected according to the result. Subsequently, OCR technology is applied to extract text and its location information from the processed image, and this information is input into a fine-tuned BERT model for sensitive entity recognition. Based on the location of the identified sensitive entities, the system performs an irreversible masking operation to protect privacy. Finally, the desensitized image can be safely stored or output. The entire process runs efficiently on resource-constrained mobile devices, not only ensuring the convenience and real-time nature of the operation, but also strictly complying with the requirements of privacy protection regulations, providing strong data security protection.
[0017] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic flowchart of the method for automatically masking sensitive clinical trial data based on a mobile terminal provided by an embodiment of the present invention; Figure 2 It is a schematic block diagram of the system for automatically masking sensitive clinical trial data based on a mobile terminal provided by an embodiment of the present invention; Figure 3 It is a schematic block diagram of the computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0022] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0023] It should be further understood that the term " / and" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0024] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the method for automatically masking sensitive data in clinical trials based on a mobile terminal provided by an embodiment of the present invention. The method for automatically masking sensitive data in clinical trials based on a mobile terminal is applied to a terminal, which interacts with a server for data. By integrating a variety of advanced image processing and deep learning technologies, efficient and accurate masking of sensitive data in clinical trials is achieved on a mobile device to meet strict privacy protection requirements. First, an image containing clinical trial data is captured by the mobile device camera, and machine learning is used to automatically detect the document border for cropping to obtain an initial image. Then, a trained convolutional neural network model is used to distinguish between paper documents and screenshots, and the best processing strategy is selected according to the classification result. Subsequently, preprocessing steps such as denoising and contrast enhancement and OCR technology are used to extract the text content and location information. The extracted results are input into a fine-tuned MobileBERT model, and the locations of sensitive entities are identified by combining context semantics, and then irreversible masking operations are performed on these locations to ensure data security. The entire process uses a series of lightweight models (such as PP-OCRv4 and MobileBERT) and optimization measures (such as exporting the model to the ONNX format and quantization, etc.) to ensure high operational convenience and real-time performance on mobile devices with limited resources, while ensuring data security and privacy.
[0025] Figure 1 FIG. 1 is a flow chart of a method for automatically masking sensitive clinical trial data based on a mobile terminal provided by an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110 to S160.
[0026] S110, acquiring an image containing clinical trial data through a mobile device camera, automatically detecting a document border and cropping a valid area to obtain an initial image.
[0027] In this embodiment, the initial image refers to a clear document image that is obtained through a mobile device camera and has been automatically border detected and cropped and optimized, and only contains clinical trial data and no unnecessary background.
[0028] In one embodiment, the above-mentioned step S110 may include steps S111 to S113.
[0029] S111. Acquire an image containing clinical trial data through a camera of a mobile device.
[0030] In this embodiment, the camera of the mobile device is used to capture images containing clinical trial information, including but not limited to paper medical records, test orders, electronic medical record interfaces, or any screenshots showing clinical trial data.
[0031] S112: Apply machine learning technology to automatically identify the document edge of the image and generate a bounding box.
[0032] In this embodiment, advanced machine learning technology is used to automatically analyze and identify the document edges in the image and generate an accurate bounding box. This process relies on the algorithm's intelligent recognition of geometric shapes and feature points in the image to ensure that the bounding box accurately surrounds the desired document area.
[0033] S113 , cropping the image based on the determined bounding box coordinates, and extracting and optimizing the cropped area using OpenCV tools to obtain an initial image.
[0034] In this embodiment, the original image is cropped according to the bounding box coordinates determined in the previous step, the region is extracted using OpenCV tools, and further optimized. This step is intended to remove unnecessary background interference, ensure that the cropped image only contains the target document content, and that the image quality meets the requirements of subsequent processing, and ultimately obtain a clear and accurate initial image.
[0035] The whole process effectively combines image acquisition, automatic border detection and cropping optimization technology to provide high-quality input images for subsequent data processing.
[0036] Specifically, images containing clinical trial data (such as medical records, test reports, etc.) are obtained through the camera of a mobile device and preliminarily processed. To ensure that the objects being processed are limited to the effective document area, this module integrates an automatic border detection algorithm that can intelligently identify the document boundaries and crop the target area. At the same time, it supports CRC to manually adjust the cropping frame to adapt to complex scenarios (such as cluttered backgrounds or tilted documents). The cropped image is used as the input for the subsequent module, which can significantly reduce the interference of irrelevant backgrounds and improve the processing efficiency and accuracy.
[0037] Collect the original images containing clinical trial data and extract the effective document or screen area through automatic border detection and cropping techniques. It integrates functions such as mobile device camera shooting, album picture import, automatic border detection, and user interaction adjustment to ensure that the images provided to the subsequent module are the precisely cropped target content, thereby improving the overall processing efficiency of the system and the accuracy of sensitive data recognition.
[0038] Use the native camera API of the mobile device (AVFoundation framework for iOS platform and CameraX library for Android platform) to call the camera to achieve real-time preview and high-resolution shooting. The system provides auxiliary functions such as autofocus, light detection tips, and image stabilization guidance to ensure the quality of the captured images.
[0039] Through the operating system's photo album access API (Photos framework for iOS and MediaStore API for Android), it is allowed to select existing pictures from the photo library. The imported pictures need to undergo format checks and resolution verification to ensure compliance with the requirements of subsequent processing.
[0040] Use the VNRectangleObservation API in the Apple Vision framework to analyze the geometric features of the image based on a deep learning model, accurately identify the four corner points of paper documents or computer screens, and form a rectangular bounding box. The module performs post-processing on the output results, including boundary smoothing, tilt correction, and false detection filtering, to improve the robustness and accuracy of the detection.
[0041] Adopt the Document Scanner API of Google ML Kit, which is designed specifically for document scanning, to quickly detect the edges and return the coordinates. Parameter adjustments are made according to the characteristics of clinical trial documents to optimize the detection performance.
[0042] After automatic border detection, the bounding box is displayed on the screen, and the user can fine-tune the border through touch operations to ensure that the cropping area completely covers the target content and there is no irrelevant background. The adjustment interface is intuitive and easy to use, supports real-time preview effects, and adopts lightweight image rendering technology to adapt to the performance limitations of mobile devices.
[0043] Based on the automatically or manually adjusted border coordinates, the module performs a cropping operation to generate a sub-image containing only the document or screen content. Matrix transformation and region extraction are carried out using OpenCV to ensure that the quality of the cropped image meets the requirements of subsequent modules. Finally, the cropped image is stored in JPEG format for use in subsequent processes.
[0044] S120. Use a convolutional neural network model trained by transfer learning to classify the initial image to distinguish between paper documents and computer screen screenshots, and select corresponding processing strategies according to the classification results.
[0045] In this embodiment, the processing strategy refers to selecting a specific optimization algorithm suitable for paper documents or computer screen screenshots for subsequent image processing according to the classification results of the convolutional neural network model on the initial image.
[0046] Identify the images output in the previous stage to determine whether they are paper document images (such as medical records, test reports), computer screen screenshots (such as electronic medical record interfaces), or other types of images (such as backgrounds or device surfaces). This process provides a basis for subsequent image processing modules to apply specific optimization algorithms according to the image type. This module uses transfer learning technology to train a convolutional neural network model to build an efficient three-classification model, and combines ONNX format conversion and quantization compression technology with the ONNX Runtime inference framework to ensure efficient deployment and operation on mobile devices. Based on the classification results, the system can select the best processing strategy to ensure the stability and accuracy of the system on different image types.
[0047] In one embodiment, the training process of the above-mentioned convolutional neural network model includes: Collect various types of desensitized images from historical records; Use LabelStudio to classify and label the collected images, and divide the labeled images to obtain a sample set; Select YOLOv11m as the base model, and train the base model through transfer learning combined with the sample set; Export the trained model in ONNX format.
[0048] In this embodiment, samples are selected from the desensitized historical images accumulated by the system in the past, including paper documents (such as handwritten or printed medical records, test reports), computer screen screenshots (such as electronic medical record interfaces), and other types of images (such as backgrounds, device surfaces). These data are from the actual clinical trial environment, ensuring diversity and representativeness.
[0049] The open-source tool LabelStudio is used for data annotation of image classification tasks. Through LabelStudio, annotators classify and label images, marking them into three categories: "PAPER", "SCREEN", and "OTHER", providing an intuitive interface and an efficient annotation process.
[0050] The labeled dataset is divided into a training set (70%), a validation set (20%), and a test set (10%) proportionally, with a total of approximately 5000 images, covering various shooting conditions such as lighting changes and background complexity.
[0051] To improve the generalization ability of the model, data augmentation techniques are used, including operations such as random cropping, rotation, brightness adjustment, contrast enhancement, noise addition, and flipping, to simulate non-ideal shooting conditions in the real world.
[0052] The YOLOv11m (a medium-sized YOLOv11 model) is selected as the base model. YOLOv11m is a medium-sized model in the YOLOv11 series, with approximately 10 million parameters and a model size of about 22.4MB, balancing performance and computational efficiency, and is suitable for classification tasks on mobile devices. Table 1 shows the parameters for model training.
[0053] Table 1 Model training parameters Parameter Item Parameter Value Number of Training Epochs 100, Early Stopping Mechanism Batch Size 16 Input Image Size 640 * 640 Optimizer Adam Learning Rate 0.001 Training Device Nvidia RTX4090D After training is completed, the model in PyTorch format is exported to ONNX format to ensure cross-platform inference compatibility, and is verified to ensure output consistency (error < 1e-5).
[0054] Post-training quantization technology is used to quantize floating-point parameters (FP32) into integer type (INT8), reducing the model size to 1 / 3 to 1 / 4 of the original size, increasing the inference speed by 2 - 3 times, while keeping the accuracy loss within 1%.
[0055] ONNX Runtime is used as the mobile inference engine to achieve efficient inference on iOS and Android platforms and support hardware acceleration (such as Core ML for iOS and NNAPI for Android).
[0056] After preprocessing the input image, it is fed into the quantized model, and the classification result and its confidence are obtained through the Softmax function. If the confidence is lower than the set threshold (0.7), the image is marked as "OTHER", prompting to reshoot or confirm.
[0057] The multi-threaded optimization and memory management functions of ONNX Runtime help reduce resource occupancy and further improve efficiency.
[0058] Based on the YOLOv11m image classification model, the model is trained through transfer learning technology to accurately distinguish the source types of input images, that is, photographed paper documents or computer screen images. For different types of images, the system adopts differentiated processing strategies: paper document images need to focus on dealing with the diversity of handwritten text and printed text, while computer screen images need to additionally handle interferences such as screen reflections and moiré patterns. The image classification results provide guidance for subsequent processing and OCR recognition, ensuring the pertinence and efficiency of the processing process.
[0059] S130. Process the initial image according to the processing strategy, and extract the text content and location information in the processed image through OCR technology to obtain an extraction result.
[0060] In this embodiment, the extraction result refers to the text content and its location information accurately recognized and extracted from the processed image through OCR technology.
[0061] According to the results of image classification, the cropped images from the image acquisition module are processed specifically, and the text content and its location information are extracted therefrom. For the different characteristics of paper documents and computer screen images, the image quality is optimized and specific interference factors are eliminated to ensure high-quality input for the subsequent sensitive information recognition module. In particular, this module performs excellently in the clinical trial scenario, can effectively process complex image characteristics (such as handwritten text, screen moiré patterns, table structures), and at the same time considers the resource limitations of mobile devices.
[0062] In one embodiment, the above step S130 may include steps S131 to S133.
[0063] S131. Process the initial image belonging to the paper document image through denoising, contrast enhancement and binarization, remove the moiré pattern in the initial image belonging to the computer screen image, and perform denoising, contrast enhancement and binarization to obtain a processing result.
[0064] In this embodiment, the processing result refers to an optimized image obtained after denoising, contrast enhancement, binarization processing of paper document images and computer screen images and moiré pattern elimination for screen images.
[0065] Specifically, due to the shooting environment or the characteristics of the document itself, it may be difficult to recognize the text in paper document images. This function designs a special preprocessing process to significantly improve the image quality, including: Denoising: Use a method combining Gaussian blur and median filtering to remove noise points, and adaptive threshold filtering further optimizes low-quality images.
[0066] Contrast Enhancement: The CLAHE algorithm is used to locally enhance the image contrast, highlighting the text area while avoiding noise amplification.
[0067] Binarization: The Otsu adaptive threshold binarization algorithm is applied to convert the image into a black-and-white binary image, improving the contrast between the text and the background. For handwritten documents, the Niblack method is also combined for local binarization processing.
[0068] To overcome the moiré pattern problem in computer screen images, this embodiment adopts the Fast Fourier Transform (FFT) technology. By performing frequency domain analysis, periodic high-frequency components are identified and removed, while the information in the text area is retained. This step significantly improves the signal-to-noise ratio of the processed image and reduces the OCR error rate.
[0069] S132. Identify horizontal and vertical lines in the processing result to obtain the recognition result.
[0070] In this embodiment, the recognition result refers to the horizontal and vertical table lines identified on the optimized image through techniques such as the Hough transform, used to determine the table boundaries and cell positions, providing a basis for subsequent text block partitioning and structured output.
[0071] Specifically, clinical trial documents often contain table structures, so accurate identification of table lines is crucial for text block partitioning. This step first uses the Hough transform to detect horizontal and vertical lines, and then uses the DBSCAN clustering algorithm to merge adjacent or overlapping line segments to form continuous table boundaries and cell positions, providing support for subsequent text block positioning.
[0072] S133. Extract the text content and position information in the processing result through OCR technology according to the recognition result to obtain the extraction result.
[0073] Specifically, according to the recognition result, the PP-OCRv4 model is used to identify and sort the text content and position information in the processing result on ONNX Runtime to obtain the extraction result.
[0074] In this embodiment, in the OCR stage, the open-source PaddleOCR system is selected for this embodiment's method, and the PP-OCRv4 model (15.8MB) suitable for mobile inference is adopted. After conversion and optimization, this model can run efficiently on ONNX Runtime. Finally, it is sorted and output according to the position coordinates of the text blocks to the next module.
[0075] Process the cropped image to improve the accuracy of text recognition. The processing steps include operations such as image denoising, binarization, contrast enhancement, etc. For computer screen images, a dedicated moiré removal algorithm is also included to eliminate interference caused by screen display. Subsequently, the module uses efficient OCR (Optical Character Recognition) technology to extract text from the processed image, obtaining the text content in the image and its corresponding position coordinates (such as bounding box information). The computing efficiency is optimized during mobile operation to ensure fast response and low resource consumption.
[0076] Through a series of finely designed technical means, the accuracy and efficiency of extracting text from images from different sources have been effectively improved, especially showing excellent performance in fields such as clinical trials that require high precision.
[0077] S140. Input the extraction result into the fine-tuned BERT model, and combine the context semantics to recognize sensitive entities to obtain the position information of the sensitive entities.
[0078] In this embodiment, the position information of the sensitive entities refers to inputting the extraction result into the fine-tuned BERT model, using its understanding of the context semantics to recognize and locate the sensitive entities in the text, so as to determine the specific positions of these entities in the document.
[0079] After processing the extraction result through segmentation and encoding, input it into the fine-tuned BERT model, and combine the context semantics to recognize sensitive entities to obtain the position information of the sensitive entities.
[0080] In one embodiment, the training process of the fine-tuned MobileBERT model includes: Select a lightweight MobileBERT model; Construct and label a clinical trial text dataset; Enhance the clinical trial text dataset through strategies such as synonym replacement, entity forgery, noise injection, and format variation to obtain a rich result; Use the rich result to fine-tune the MobileBERT model for specific sensitive entity recognition tasks based on the sequence labeling task using the cross-entropy loss function and the AdamW optimizer to obtain the fine-tuned BERT model; Export the fine-tuned BERT model in ONNX format.
[0081] Based on a small-scale BERT pre-trained language model suitable for mobile deployment, through fine-tuning for the task of identifying sensitive entities in clinical trials, accurate identification of sensitive information in text is achieved. Sensitive information includes, but is not limited to, personally identifiable information (PII) such as patient names, ID numbers, phone numbers, addresses, medical record numbers, etc. The fine-tuned BERT model can combine context semantics, accurately distinguish sensitive entities from non-sensitive content, and output the location information of sensitive entities. When designing the model, the computing power limitations of mobile devices are fully considered, and model compression and quantization techniques are used to ensure efficient operation.
[0082] In this embodiment, sensitive information (such as personally identifiable information like patient names, ID numbers, phone numbers, etc., abbreviated as PII) is accurately identified in the text output from the OCR recognition module. This module uses a small-scale BERT pre-trained language model suitable for mobile deployment, and through fine-tuning for the task of identifying sensitive entities in clinical trial scenarios, efficient and accurate identification in resource-constrained environments is achieved. At the same time, mobile inference technology is combined to ensure real-time and accurate detection of sensitive information on mobile devices. This step provides key input for subsequent text block positioning and masking, directly affecting the accuracy and compliance of data de-sensitization.
[0083] Based on the labeled dataset of clinical trial scenarios, the small-scale BERT model is fine-tuned to optimize its performance in the task of identifying sensitive entities.
[0084] The small BERT model (MobileBERT) is selected as the base model. MobileBERT is a lightweight variant of BERT, optimized through knowledge distillation and intra-layer decomposition techniques, retaining the semantic understanding ability of the Transformer encoder architecture and adapting to the computing and storage limitations of mobile devices. Table 2 shows the basic parameters of MobileBERT used in this embodiment: Table 2. Basic Parameters of the Selected MobileBERT Model in This Embodiment Parameter Item Parameter Value Number of Transformer Layers 12 Number of Parameters 25.3M Number of Attention Heads 8 Hidden Size 512 Maximum Sequence Length 512 tokens Word Embedding Dimension 512 Pre-trained Weights mobilebert-uncased, Supporting Multilingual Texts Such as Chinese and English Construct a high-quality labeled dataset that covers diverse texts and sensitive entities in clinical trial scenarios to enhance the robustness of the model to OCR noise and complex expressions.
[0085] Collect OCR-extracted text data from clinical trial scenarios, and the sources include paper medical records, laboratory test reports, informed consent forms, electronic medical record interfaces, etc. The data covers sensitive entity types as shown in Table 3 below.
[0086] Table 3. Sensitive Entity Types Serial Number Chinese Name English Abbreviation 1 Name NAME 2 Contact Number PHONE 3 Address ADDRESS 4 ID Card Number ID_NUMBER 5 Work Unit WORKPLACE 6 Medical Record Number MEDICAL_RECORD_NUMBER 7 Outpatient Number OUTPATIENT_NUMBER 8 Patient ID PATIENT_ID 9 Hospitalization Number HOSPITALIZATION_NUMBER 10 Examination Number EXAMINATION_NUMBER
[0087] Approximately 5000 text fragments (about 1 million tokens), including handwritten text transcriptions, mixed Chinese and English texts, and entities with irregular formats.
[0088] Use the open-source LabelStudio platform and organize professionals for entity-level annotation. Adopt BIO (Begin, Inside, Outside) annotation to assign labels to each token (e.g., "B-PER" indicates the start of a name, "I-PER" indicates inside a name, and "O" indicates a non-entity). For 10 main types of sensitive entities, an additional "Other" category (O) is used to handle non-sensitive texts. Through multiple rounds of review in LabelStudio, calculate the annotation consistency (Cohen's Kappa coefficient ≥ 0.85) to ensure annotation accuracy.
[0089] The following strategies are adopted for data augmentation: Synonym replacement: Based on medical domain dictionaries (such as ICD-10, SNOMED CT), replace synonymous terms (e.g., replace "patient name" with "patient's first name").
[0090] Entity forgery: Generate virtual entities (such as a random name "Li Ming" and an ID number "××××××") to increase entity diversity.
[0091] Noise injection: Simulate OCR misrecognition and introduce 10%-20% noise (such as spelling mistakes like changing "Zhang Wei" to "Zhang Wei", word segmentation like changing "×××××××××××" to "××× ×××× ×××× ", and punctuation omission).
[0092] Format variation: Generate entity variations through regular expressions (such as a phone number "×××-××××-××××" or "×××××××××××").
[0093] The dataset is divided according to the following ratio: Training set: 70% (3500 pieces), validation set: 20% (1000 pieces), test set: 10% (500 pieces).
[0094] The validation set is used for hyperparameter tuning, and the test set is used to evaluate the model performance (target F1 score ≥ 90%).
[0095] Sensitive entity recognition is a sequence annotation task (Token Classification), predicting BIO labels (such as "B-PER", "I-PER", "O") for each token; Load the pre-trained weights of mobilebert-uncased, retain the encoder parameters, and initialize the classification head.
[0096] Freeze the first 8 layers (about 65% of the number of parameters), fine-tune the last 4 layers and the classification head to reduce computational overhead and prevent overfitting.
[0097] The loss function is Cross-Entropy Loss, combined with label smoothing (smoothing factor 0.1) to enhance robustness.
[0098] The optimizer is AdamW, with an initial learning rate of 2e-5 and a weight decay of 1e-2.
[0099] The learning rate schedule is Linear Schedule with Warmup. There is a warmup for the first 1000 steps (about 10% of the iterations), and the learning rate linearly increases from 0 to 2e-5.
[0100] The number of training epochs is 15, with an early stopping mechanism (terminate when the F1 on the validation set does not improve by <0.005 for 5 consecutive epochs).
[0101] The batch size is 8, combined with gradient accumulation (equivalent Batch Size = 32) to simulate large-batch training.
[0102] Use an 8-GPU NVIDIA RTX 4090D server (24 * 8 GB of video memory) for training.
[0103] Since the number of parameters of the selected model is not large, the model size is about 60M. After testing, its inference speed meets the requirements, so the model is directly exported to the ONNX format and the computational graph is optimized (node fusion, constant folding).
[0104] In this embodiment, the output text blocks are directly separated by spaces. These text blocks may contain various types of sensitive information, such as names, phone numbers, etc. This step ensures that subsequent processing can analyze the text in the most basic form.
[0105] Since the maximum sequence length of the model is limited to 512 tokens, the long text output by OCR needs to be split into multiple short text paragraphs that do not exceed this limit. The purpose of this is to avoid exceeding the maximum length limit supported by the model during processing, thus ensuring that the model can run properly and make accurate predictions.
[0106] Use the MobileBERT tokenizer (mobilebert-uncased) to encode the segmented text, converting it into token IDs, attention masks, and token type IDs. In addition, special tokens such as [CLS] and [SEP] are added to facilitate the model's understanding of the start and end of sentences. This step is necessary because it converts the original text into a form that the model can understand and process.
[0107] The encoded sequence obtained after word segmentation is fed into a pre-trained model loaded by ONNX Runtime. The ONNX format is mainly adopted to improve the inference speed and efficiency, enabling the model to run efficiently in resource-constrained environments.
[0108] The model predicts the probability distribution of BIO tags for each token (a total of 11 dimensions, corresponding to the Softmax output). This means that for each token, the model gives the likelihood that it belongs to a specific entity (such as the start "B-PER" or the middle part "I-PER" of a person's name).
[0109] After obtaining the prediction results for all tokens, the next task is to merge consecutive BIO tags to form complete entities. For example, the "B-PER" and subsequent "I-PER" tags are combined to form a complete person's name entity (such as "Zhang Wei"). This process also includes generating entity-level results, that is, determining the specific text content of the entity, its type, the start and end index positions in the original text, and the confidence score.
[0110] A confidence threshold (0.9) is set, and all entities below this threshold are marked as "to be confirmed" and the clinical research coordinator (CRC) is prompted to conduct manual verification. This mechanism helps to reduce misidentifications. By manually reviewing those entities that the model is less certain about, it is expected that the misidentification rate can be reduced by approximately 5% - 10%.
[0111] Based on the text position information recognized by OCR and the sensitive entity recognition results output by the BERT model, text block localization and irreversible masking operations are performed on the sensitive information regions in the image. Text block localization includes steps such as bounding box merging and misidentification filtering to improve the accuracy of the masked region. The masking operation uses an irreversible covering method (such as black filling or mosaic processing) to ensure that sensitive information cannot be restored by technical means. The finally output desensitized image meets the compliance requirements for clinical trial data privacy protection and can be used for subsequent archiving or analysis.
[0112] Through the above steps, not only the accuracy of sensitive information recognition is improved, but also it is ensured that the system can achieve effective real-time sensitive information detection on resource-limited platforms such as mobile devices while maintaining high efficiency.
[0113] S150. According to the position information of the sensitive entity, perform an irreversible masking operation on the corresponding region in the image containing clinical trial data to obtain a desensitized image.
[0114] In this embodiment, the desensitized image refers to an image obtained by performing an irreversible masking operation on the corresponding region in the image containing clinical trial data according to the position information of the sensitive entity, with the sensitive information removed.
[0115] Specifically, use black filling and apply a 3×3 Gaussian blur to the edge or use manual masking to perform an irreversible masking operation on the corresponding region in the image containing clinical trial data according to the position information of the sensitive entity, so as to obtain the desensitized image.
[0116] In this embodiment, text block localization is the process of mapping the sensitive entities (including their entity text, type, token index, and confidence) determined in the sensitive information recognition module to the text block bounding box coordinates provided by the OCR module. The goal of this process is to generate an accurate masking region bounding box to accurately cover the sensitive information in the image. To address possible issues such as OCR text segmentation errors, same-text ambiguity, and table structure complexity, this module adopts the following two strategies: To accurately match entities with the same text content, this strategy narrows down the scope by analyzing the context information around the sensitive entity. Specifically, it considers the content of the 5 characters before and after each sensitive entity and performs an exact match in the OCR result list based on this information. Here, a string matching algorithm is used, and the most suitable match is found according to the edit distance (threshold ≤ 2). This method helps to solve the ambiguity problem that occurs when there are multiple identical sensitive entities in the document.
[0117] Table partitioning constraint: When processing a document containing a table, use the table line positions detected by the image processing module to limit the scope of the text block where the sensitive entity is located. This can ensure that the generated masking region bounding box does not cross the table lines (either horizontally or vertically), thus more accurately locating the sensitive information. This approach is particularly important for maintaining the integrity and clarity of the table structure.
[0118] Once the specific positions of the sensitive information to be masked are determined, the next step is to perform an irreversible masking operation on these regions. The masking process is divided into two modes: automatic and manual. The automatic masking function uses black filling (RGB: 0, 0, 0) to cover the regions marked as needing to be masked. To make the masking edge look more natural and smooth, a 3×3 Gaussian blur effect is applied for feathering. This can not only effectively hide the sensitive information but also ensure the overall aesthetics of the document.
[0119] Manual masking provides a user - interactive masking method that allows users to directly draw masking areas on images through a touch interface. This function includes a brush tool with a width of 5 pixels and a rectangular selection tool, enabling users to flexibly adjust the masking area according to actual needs. This is very useful for situations where the system fails to correctly identify or additional privacy protection is required.
[0120] Through the above steps, the text block positioning and automatic masking module can not only effectively protect sensitive information in the document, but also ensure that the processed document maintains good readability and visual consistency. This process comprehensively applies knowledge and technologies in multiple aspects such as OCR technology, string matching algorithms, and image processing technology, reflecting a high degree of professionalism and practicality.
[0121] S160. Store or output the desensitized image.
[0122] Store the desensitized image in the server or output it to other terminals, etc.
[0123] The method of this embodiment aims to solve the defects existing in the prior art, such as low efficiency of manual smearing, insufficient accuracy of rule recognition, and possible violation of compliance requirements in server - side processing. By integrating image processing, deep learning, and pre - trained language model technologies, the present invention can achieve efficient, accurate, and regulatory - compliant sensitive data masking on mobile devices. Its main advantages include: High - efficiency automation and high - precision recognition: Integrating automatic cropping, image classification (using convolutional neural networks), OCR recognition, and fine - tuning the BERT model for sensitive information recognition, it realizes the full - process automation from image acquisition to sensitive information masking. The F1 score of sensitive entity recognition in this system reaches 91% - 93%, which is significantly higher than the accuracy of traditional manual smearing and rule - matching methods, greatly reducing the workload of clinical research coordinators.
[0124] Mobile - end local processing ensures compliance: All processing steps are completed locally on the mobile device without transmitting the original data to an external server, meeting the requirements of regulations such as GDPR, HIPAA, and the Personal Information Protection Law that clinical trial data shall not leave the research center, thus completely avoiding the regulatory risks brought by data leakage and server - side processing.
[0125] Robustness in complex scenarios: For various complex situations in clinical trials, such as handwritten text, screen moiré patterns, low - light conditions, table structures, and mixed Chinese - English text, the present invention improves the robustness and adaptability of the system through differential pre - processing, noise sample training, and context enhancement technologies, enabling the accuracy of text recognition and sensitive entity recognition to reach over 90% and 85% respectively.
[0126] Resource-Efficient Adaptation to Mobile Devices: By adopting model compression, lightweight algorithms, and hardware acceleration technologies, the system can run quickly on resource-constrained mobile devices. The processing time for a single image is less than 3 seconds, and the memory occupancy does not exceed 100MB, solving the performance bottleneck problem on mobile devices with low computing power.
[0127] Scalability and Internationalization Support: The system not only supports the processing of mixed Chinese and English texts but also can add new sensitive entity types by updating the dataset and fine-tuning the model, and be extended to other languages in the future to meet the international clinical trial requirements and the ever-changing privacy regulation requirements.
[0128] The above-mentioned automatic masking method for sensitive clinical trial data based on mobile devices realizes the efficient and accurate masking of sensitive clinical trial data on mobile devices through the integration of a series of advanced technologies, while ensuring data security and privacy. First, the device camera is used to obtain an image and automatically detect and crop the document border to extract the effective area. Then, a convolutional neural network model trained by transfer learning is used to classify the image, and an appropriate processing strategy is selected according to the result. Subsequently, OCR technology is applied to extract the text and its location information from the processed image, and this information is input into the fine-tuned BERT model for sensitive entity recognition. Based on the identified sensitive entity locations, the system performs irreversible masking operations to protect privacy. Finally, the desensitized image can be safely stored or output. The whole process runs efficiently on resource-constrained mobile devices, not only ensuring the convenience and real-time nature of the operation but also strictly following the requirements of privacy protection regulations, providing strong data security guarantees.
[0129] Figure 2 It is a schematic block diagram of an automatic masking system 300 for sensitive clinical trial data based on mobile devices provided by an embodiment of the present invention. As Figure 2 shown, corresponding to the above automatic masking method for sensitive clinical trial data based on mobile devices, the present invention also provides an automatic masking system 300 for sensitive clinical trial data based on mobile devices. The automatic masking system 300 for sensitive clinical trial data based on mobile devices includes units for executing the above automatic masking method for sensitive clinical trial data based on mobile devices, and this device can be configured in terminals such as desktop computers, tablet computers, laptops, etc. Specifically, please refer to Figure 2 . The automatic masking system 300 for sensitive clinical trial data based on mobile devices includes an image acquisition unit 301, a classification unit 302, a processing unit 303, an identification unit 304, a masking unit 305, and a storage and output unit 306.
[0130] An image acquisition unit 301 is configured to acquire an image containing clinical trial data through a mobile device camera, automatically detect the document border and crop the valid area to obtain an initial image; a classification unit 302 is configured to classify the initial image by using a convolutional neural network model trained by transfer learning to distinguish between a paper document and a computer screen screenshot, and select a corresponding processing strategy according to the classification result; a processing unit 303 is configured to process the initial image according to the processing strategy, and extract the text content and location information in the processed image through OCR technology to obtain an extraction result; an identification unit 304 is configured to input the extraction result into a fine-tuned BERT model, and identify sensitive entities by combining context semantics to obtain the location information of the sensitive entities; a masking unit 305 is configured to perform an irreversible masking operation on the corresponding area in the image containing clinical trial data according to the location information of the sensitive entities to obtain a desensitized image; a storage and output unit 306 is configured to store or output the desensitized image.
[0131] In one embodiment, the image acquisition unit 301 includes: An image acquisition subunit is configured to acquire an image containing clinical trial data through a mobile device camera; an edge recognition subunit is configured to automatically recognize the document edge of the image by using machine learning technology to generate a bounding box; a cropping subunit is configured to crop the image based on the determined bounding box coordinates, and use the OpenCV tool to extract and optimize the cropped area to obtain an initial image.
[0132] In one embodiment, the extraction unit includes: A processing subunit is configured to process the initial image belonging to the paper document image through denoising, contrast enhancement and binarization, remove moiré patterns in the initial image belonging to the computer screen image, and perform denoising, contrast enhancement and binarization to obtain a processing result; a straight line recognition subunit is configured to recognize horizontal and vertical straight lines in the processing result to obtain a recognition result; a content extraction subunit is configured to extract the text content and location information in the processing result through OCR technology according to the recognition result to obtain an extraction result.
[0133] In one embodiment, the content extraction subunit is configured to recognize and sort the text content and location information in the processing result by using the PP-OCRv4 model on ONNX Runtime according to the recognition result to obtain an extraction result.
[0134] In one embodiment, the identification unit 304 is configured to input the extraction result into a fine-tuned BERT model after segmentation and encoding processing, and identify sensitive entities by combining context semantics to obtain the location information of the sensitive entities.
[0135] In one embodiment, the masking unit 305 is configured to perform an irreversible masking operation on the corresponding region in the image containing clinical trial data by filling it with black and applying a 3×3 Gaussian blur to the edges or by manually masking according to the position information of the sensitive entity, so as to obtain a desensitized image.
[0136] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above-mentioned mobile-based automatic masking system 300 for clinical trial sensitive data and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity of description, they will not be elaborated herein.
[0137] The above-mentioned mobile-based automatic masking system 300 for clinical trial sensitive data can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 3 shown.
[0138] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 may be a terminal. Among them, the terminal may be an electronic device with communication functions such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device.
[0139] Referring to Figure 3 , the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.
[0140] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions. When the program instructions are executed, the processor 502 can be made to execute a mobile-based method for automatically masking clinical trial sensitive data.
[0141] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0142] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be made to execute a mobile-based method for automatically masking clinical trial sensitive data.
[0143] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 3The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. Specifically, the computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0144] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the following steps: Obtain an image containing clinical trial data through the mobile device camera, automatically detect the document border and crop out the effective area to obtain an initial image; use a convolutional neural network model trained by transfer learning to classify the initial image to distinguish between paper documents and computer screen screenshots, and select corresponding processing strategies according to the classification results; process the initial image according to the processing strategy, and extract the text content and location information in the processed image through OCR technology to obtain an extraction result; input the extraction result into the fine-tuned BERT model, and identify sensitive entities by combining context semantics to obtain the location information of the sensitive entities; according to the location information of the sensitive entities, perform an irreversible masking operation on the corresponding area in the image containing clinical trial data to obtain a desensitized image; store or output the desensitized image.
[0145] Among them, the training process of the convolutional neural network model includes: Collect various types of desensitized images from historical records; use LabelStudio to classify and label the collected images, and divide the labeled images to obtain a sample set; select YOLOv11m as the basic model, and train the basic model through transfer learning in combination with the sample set; export the trained model in ONNX format.
[0146] The training process of the fine-tuned BERT model includes: Select the lightweight MobileBERT model; construct and label the clinical trial text data set; enhance the clinical trial text data set through strategies such as synonym replacement, entity forgery, noise injection, and format variation to obtain a rich result; use the rich result to fine-tune the MobileBERT model for specific sensitive entity recognition tasks based on the sequence labeling task using the cross-entropy loss function and the AdamW optimizer to obtain the fine-tuned BERT model; export the fine-tuned BERT model in ONNX format.
[0147] In an embodiment, when the processor 502 implements the step of obtaining an image containing clinical trial data through the mobile device camera, automatically detecting the document border and cropping out the effective area to obtain an initial image, the following steps are specifically implemented: Obtain an image containing clinical trial data through the camera of a mobile device; apply machine learning techniques to automatically identify the document edges of the image and generate a bounding box; crop the image based on the determined bounding box coordinates, and use OpenCV tools to extract and optimize the cropped area to obtain an initial image.
[0148] In one embodiment, when the processor 502 implements the step of processing the initial image according to the processing strategy and extracting the text content and location information in the processed image through OCR technology to obtain an extraction result, the specific implementation is as follows: Perform denoising, contrast enhancement, and binarization on the initial image belonging to a paper document image, remove moiré patterns in the initial image belonging to a computer screen image, and perform denoising, contrast enhancement, and binarization to obtain a processing result; perform horizontal and vertical line recognition on the processing result to obtain a recognition result; extract the text content and location information in the processing result through OCR technology according to the recognition result to obtain an extraction result.
[0149] In one embodiment, when the processor 502 implements the step of extracting the text content and location information in the processing result through OCR technology according to the recognition result to obtain an extraction result, the specific implementation is as follows: Use the PP-OCRv4 model to recognize and sort the text content and location information in the processing result on ONNX Runtime according to the recognition result to obtain an extraction result.
[0150] In one embodiment, when the processor 502 implements the step of inputting the extraction result into a fine-tuned BERT model and identifying sensitive entities in combination with context semantics to obtain the location information of the sensitive entities, the specific implementation is as follows: After splitting and encoding the extraction result, input it into the fine-tuned BERT model, and identify sensitive entities in combination with context semantics to obtain the location information of the sensitive entities.
[0151] In one embodiment, when the processor 502 implements the step of performing an irreversible masking operation on the corresponding area in the image containing clinical trial data according to the location information of the sensitive entities to obtain a desensitized image, the specific implementation is as follows: Use black filling and apply a 3×3 Gaussian blur to the edge or perform manual masking. According to the location information of the sensitive entities, perform an irreversible masking operation on the corresponding area in the image containing clinical trial data to obtain a desensitized image.
[0152] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit 303 (CPU). The processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0153] Those of ordinary skill in the art can understand that all or part of the processes in the methods of implementing the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0154] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the following steps: Obtain an image containing clinical trial data through a mobile device camera, automatically detect the document border and crop out the valid area to obtain an initial image; use a convolutional neural network trained by transfer learning to classify the initial image to distinguish between paper documents and computer screen screenshots, and select corresponding processing strategies according to the classification results; process the initial image according to the processing strategies, and extract the text content and location information in the processed image through OCR technology to obtain an extraction result; input the extraction result into a fine-tuned BERT model, and identify sensitive entities in combination with context semantics to obtain the location information of the sensitive entities; according to the location information of the sensitive entities, perform an irreversible masking operation on the corresponding area in the image containing clinical trial data to obtain a desensitized image; store or output the desensitized image.
[0155] Among them, the training process of the convolutional neural network model includes: Collect various types of desensitized images from the historical records; classify and label the collected images using LabelStudio, and divide the labeled images to obtain a sample set; select YOLOv11m as the base model, and train the base model through transfer learning combined with the sample set; export the trained model in ONNX format.
[0156] The training process of the fine-tuned BERT model includes: Select a lightweight MobileBERT model; construct and label a clinical trial text dataset; enhance the clinical trial text dataset through synonym replacement, entity forgery, noise injection, and format variation strategies to obtain a rich result; use the rich result to fine-tune the MobileBERT model for specific sensitive entity recognition tasks based on a sequence labeling task using a cross-entropy loss function and an AdamW optimizer to obtain a fine-tuned BERT model; export the fine-tuned BERT model in ONNX format.
[0157] In one embodiment, when the processor executes the computer program to implement the step of obtaining an image containing clinical trial data through the mobile device camera, automatically detecting the document border and cropping out the effective area to obtain an initial image, the specific implementation steps are as follows: Obtain an image containing clinical trial data through the mobile device camera; apply machine learning techniques to automatically identify the document edge of the image and generate a bounding box; crop the image based on the determined bounding box coordinates, and use OpenCV tools to extract and optimize the cropped area to obtain an initial image.
[0158] In one embodiment, when the processor executes the computer program to implement the step of processing the initial image according to the processing strategy and extracting the text content and location information in the processed image through OCR technology to obtain an extraction result, the specific implementation steps are as follows: Perform denoising, contrast enhancement, and binarization on the initial image belonging to a paper document image, remove moiré patterns in the initial image belonging to a computer screen image, and perform denoising, contrast enhancement, and binarization to obtain a processing result; perform horizontal and vertical line recognition on the processing result to obtain a recognition result; extract the text content and location information in the processing result through OCR technology according to the recognition result to obtain an extraction result.
[0159] In one embodiment, when the processor executes the computer program to implement the step of extracting the text content and location information in the processing result through OCR technology according to the recognition result to obtain an extraction result, the specific implementation steps are as follows: Based on the recognition result, use the PP-OCRv4 model to recognize and sort the text content and location information in the processing result on ONNX Runtime to obtain the extraction result.
[0160] In one embodiment, when the processor executes the computer program to implement the step of inputting the extraction result into the fine-tuned BERT model and identifying sensitive entities in combination with the context semantics to obtain the location information of the sensitive entities, the specific implementation is as follows: After splitting and encoding the extraction result, input it into the fine-tuned BERT model, and identify sensitive entities in combination with the context semantics to obtain the location information of the sensitive entities.
[0161] In one embodiment, when the processor executes the computer program to implement the step of performing an irreversible masking operation on the corresponding region in the image containing clinical trial data according to the location information of the sensitive entities to obtain the desensitized image, the specific implementation is as follows: Use black filling and apply a 3×3 Gaussian blur to the edge or use manual masking to perform an irreversible masking operation on the corresponding region in the image containing clinical trial data according to the location information of the sensitive entities to obtain the desensitized image.
[0162] The storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.
[0163] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0164] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0165] The steps in the method of the embodiments of the present invention can be adjusted in sequence, combined, and deleted according to actual needs. The units in the device of the embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing unit 303, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0166] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention.
[0167] As mentioned above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for automatically masking sensitive data in clinical trials based on a mobile device, characterized in that Including: Obtain an image containing clinical trial data through the camera of a mobile device, automatically detect the document border and crop out the valid area to obtain an initial image; Use a convolutional neural network model trained by transfer learning to classify the initial image to distinguish between paper documents and computer screen screenshots, and select corresponding processing strategies according to the classification results; Process the initial image according to the processing strategy, and extract the text content and location information in the processed image through OCR technology to obtain an extraction result; Input the extraction result into a fine-tuned BERT model to identify sensitive entities in combination with context semantics to obtain the location information of the sensitive entities; According to the location information of the sensitive entities, perform an irreversible masking operation on the corresponding area in the image containing clinical trial data to obtain a desensitized image; Store or output the desensitized image.
2. The automatic masking method for sensitive data in clinical trials based on a mobile device according to claim 1, wherein The step of obtaining an image containing clinical trial data through the camera of a mobile device, automatically detecting the document border and cropping out the valid area to obtain an initial image includes: Obtain an image containing clinical trial data through the camera of a mobile device; Apply machine learning technology to automatically identify the document edge of the image and generate a bounding box; Crop the image based on the determined bounding box coordinates, and use OpenCV tools to extract and optimize the cropped area to obtain an initial image.
3. The method for automatically masking sensitive data in clinical trials based on a mobile device according to claim 1, wherein The training process of the convolutional neural network model includes: Collect various types of desensitized images from historical records; Use LabelStudio to classify and label the collected images, and divide the labeled images to obtain a sample set; Select YOLOv11m as the base model, and train the base model through transfer learning in combination with the sample set; Export the trained model in ONNX format.
4. The method for automatically masking sensitive data in clinical trials based on a mobile device according to claim 1, wherein The step of processing the initial image according to the processing strategy, and extracting the text content and location information in the processed image through OCR technology to obtain an extraction result includes: Process the initial image belonging to a paper document image through denoising, contrast enhancement and binarization, remove moiré patterns in the initial image belonging to a computer screen image, and perform denoising, contrast enhancement and binarization to obtain a processing result; Perform horizontal and vertical line recognition on the processing result to obtain a recognition result; Extract the text content and location information in the processing result through OCR technology according to the recognition result to obtain an extraction result.
5. The method for automatically masking sensitive data in clinical trials based on a mobile device according to claim 4, wherein The step of extracting the text content and location information in the processing result through OCR technology according to the recognition result to obtain an extraction result includes: Use the PP-OCRv4 model to identify and sort the text content and location information in the processing result on ONNX Runtime according to the recognition result to obtain an extraction result.
6. The method for automatically masking sensitive data in clinical trials based on a mobile device according to claim 1, wherein The training process of the fine-tuned BERT model includes: Select a lightweight MobileBERT model; Construct and label a clinical trial text dataset; Enhance the clinical trial text dataset through synonym replacement, entity forgery, noise injection, and format variation strategies to obtain rich results; Use the rich results to fine-tune the MobileBERT model for specific sensitive entity recognition tasks based on a sequence labeling task using a cross-entropy loss function and an AdamW optimizer to obtain a fine-tuned BERT model; Export the fine-tuned BERT model in ONNX format.
7. The method for automatically masking sensitive data in clinical trials based on a mobile device according to claim 1, wherein Input the extraction result into the fine-tuned BERT model, and identify sensitive entities by combining context semantics to obtain the location information of sensitive entities, including: After processing the extraction result through segmentation and encoding, input it into the fine-tuned BERT model, and identify sensitive entities by combining context semantics to obtain the location information of sensitive entities.
8. The method for automatically masking sensitive data in clinical trials based on a mobile device according to claim 1, wherein According to the location information of the sensitive entities, perform an irreversible masking operation on the corresponding region in the image containing clinical trial data to obtain a de-identified image, including: Perform an irreversible masking operation on the corresponding region in the image containing clinical trial data by using black filling and applying a 3×3 Gaussian blur to the edge or by using manual masking according to the location information of the sensitive entities to obtain a de-identified image.
9. The mobile device-based automatic masking system for clinical trial sensitive data is characterized in that Include: An image acquisition unit for acquiring an image containing clinical trial data through a mobile device camera, automatically detecting the document border and cropping out the valid region to obtain an initial image; A classification unit for classifying the initial image using a convolutional neural network trained by transfer learning to distinguish between paper documents and computer screen screenshots, and selecting corresponding processing strategies according to the classification results; A processing unit for processing the initial image according to the processing strategy, and extracting the text content and location information in the processed image through OCR technology to obtain an extraction result; An identification unit for inputting the extraction result into the fine-tuned BERT model, and identifying sensitive entities by combining context semantics to obtain the location information of sensitive entities; A masking unit for performing an irreversible masking operation on the corresponding region in the image containing clinical trial data according to the location information of the sensitive entities to obtain a de-identified image; A storage and output unit for storing or outputting the de-identified image.
10. A computer device, characterized in that, The computer device includes a memory and a processor, and a computer program is stored on the memory. When the processor executes the computer program, the method described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Medical sensitive data acquisition and desensitization method
CN116756750A
Image processing method and device, storage medium and computer equipment
CN117036179A
Image sensitive information protection method, server and storage medium
CN119583724A
Text keyword desensitization method, system and equipment based on large model and medium
CN119622809A
Substation main wiring diagram intelligent identification method and system based on machine learning
CN119693955A
Cited By
Unstructured data feature multimode extraction and identification method
CN121705979A