An ophthalmology clinical medical data management and sharing system

Through modular design and intelligent de-identification processing, combined with encrypted storage and automated compliance evaluation, the balance between privacy protection and data effectiveness in ophthalmic data management and sharing systems is solved, and efficient and secure data sharing is achieved.

CN120068156BActive Publication Date: 2025-07-11NANJING PEIYUE TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510488018.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-11
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing ophthalmic data management and sharing systems are difficult to balance privacy protection and data effectiveness during the de-identification process, and lack an automated compliance review mechanism, which affects the integrity and accuracy of data.

Method used

It adopts a modular design, including data extraction module, de-identification module, privacy protection module, compliance review module and data sharing real-time monitoring module. Through intelligent de-identification processing, encrypted storage and automated compliance evaluation, it ensures privacy protection and compliance of data during the sharing process.

Benefits of technology

It realizes the effectiveness and compliance of data while ensuring privacy protection, improves the security and efficiency of the data sharing process, reduces the risks of manual intervention and potential leakage, and supports clinical research and academic sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068156B_ABST
    Figure CN120068156B_ABST
Patent Text Reader

Abstract

The present invention discloses an ophthalmology clinical medical data management and sharing system, which relates to the technical field of data sharing. Through modular design, the system effectively solves many deficiencies of traditional medical data management systems in aspects such as privacy protection, data de-identification, and compliance review. Each module complements each other to ensure that privacy protection and compliance control can be achieved throughout the entire process of data collection, de-identification, and sharing. In the data extraction module, the system can automatically extract data from the ophthalmology clinical database, perform preprocessing and standardization, enabling the subsequent de-identification and data sharing processes to proceed smoothly. In addition, the de-identification module further ensures the privacy protection of patients' sensitive information through intelligent de-identification processing and encrypted storage technology, making the shared data meet the privacy protection requirements while retaining sufficient valid information for clinical use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data sharing, and particularly to an ophthalmic clinical medical data management and sharing system. Background Art

[0002] With the rapid development of digital medicine, the management and sharing of medical data have become an important part of the modern medical system. There are huge demands and challenges, especially in the aspects of medical data privacy protection and cross-institutional data sharing. Globally, medical data is gradually shifting from traditional paper records to diverse digital data such as electronic health records (EHRs), medical images, and genomic data. These data contain a large amount of sensitive information, such as patients' personal identities, health conditions, disease histories, treatment plans, etc. How to share data while ensuring privacy has become one of the key issues in modern medical management systems. Especially in the field of ophthalmology, the types of medical data involved are diverse, including fundus images, OCT scans, visual acuity test results, clinical notes, etc. These data not only need to meet the requirements of patient privacy protection but also ensure their effectiveness in academic research and clinical decision-making.

[0003] Current ophthalmic data management and sharing systems face many problems. Especially in the process of de-identification and data sharing, it is difficult to achieve a balance between privacy protection and data effectiveness. Existing technologies mostly rely on manual processing for the de-identification process. Although it can remove patients' direct identity information, it often sacrifices the effectiveness of the data, affecting subsequent academic research or clinical decision support. For example, when removing patients' names and hospital watermarks from ophthalmic image data, some medical features may be lost; when de-identifying clinical data, it may affect the personalized medical data records of patients, making the data no longer complete and accurate during analysis. In addition, there are also inconsistencies in the current compliance review process. Many systems lack an automated review mechanism and it is difficult to conduct real-time compliance assessments on different types of medical data according to sharing rules requirements, such as GDPR and HIPAA. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides an ophthalmic clinical medical data management and sharing system, which solves the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: It includes a data extraction module, a de-identification module, a privacy protection module, a compliance review module, and a data sharing real-time monitoring module;

[0006] The data extraction module extracts ophthalmic clinical data by constructing a data sharing platform and accessing an ophthalmic clinical database, and performs preprocessing to obtain standardized ophthalmic clinical data;

[0007] The de-identification module de-identifies the standardized ophthalmic clinical data through intelligent de-identification processing, obtains un-identified data items, extracts a shared data set based on the un-identified data items, and stores it using an encryption algorithm.

[0008] The privacy protection module constructs a privacy protection degree algorithm formula, calculates and outputs the privacy protection degree Ppr, sets a privacy protection threshold P, and conducts a preliminary comparison and evaluation with the privacy protection degree Ppr, and triggers a compliance review based on the preliminary comparison and evaluation.

[0009] The compliance review module constructs a compliance algorithm formula and a data quality algorithm formula, calculates and outputs a compliance score Cco and a data validity score Eva, and analyzes the compliance and validity in the un-identified data items after de-identification respectively.

[0010] The data sharing real-time monitoring module calculates and outputs a data sharing index Esha through the compliance score Cco and the data validity score Eva, sets a sharing threshold Tth, and conducts a sharing evaluation with the data sharing index Esha to judge the current abnormal situation of data de-identification.

[0011] Preferably, the data extraction module includes a data extraction unit and a data processing unit.

[0012] The data extraction unit constructs a data sharing platform, sets an API application program interface, accesses the ophthalmic clinical database, and extracts the ophthalmic clinical data sent by the user in the shared request after the user sends the shared request, and transmits it to the data sharing platform.

[0013] The ophthalmic clinical data includes image data and text data, where the text data includes clinical data, note data, and genetic information.

[0014] The data processing unit preprocesses the extracted ophthalmic clinical data, and the preprocessing includes image data processing, text data processing, data cleaning, and standardization, to obtain standardized ophthalmic clinical data.

[0015] The image data processing standardizes fundus images and retinal scans using computer vision technology to remove redundant information.

[0016] The text data processing processes notes and medical records using natural language processing NLP technology to identify sensitive information therein.

[0017] The data cleaning and standardization denoise and standardize all ophthalmic clinical data and unify the data format.

[0018] Preferably, the de-identification module includes an intelligent de-identification processing unit, a shared parameter extraction unit, and an encrypted storage unit;

[0019] The intelligent de-identification processing unit performs de-identification processing on the standardized ophthalmic clinical data through setting intelligent de-identification processing, removes sensitive information in the standardized ophthalmic clinical data, and obtains unlabeled data items. The intelligent de-identification processing includes image data de-identification processing and text de-identification processing;

[0020] The sensitive information includes privacy areas, patient information, sensitive words, and genetic information;

[0021] The image data de-identification processing locates the privacy area in the image by using a target detection algorithm, performs seamless replacement on the privacy area by using an image repair algorithm, and simultaneously uses edge detection Canny, SIFT, and SURF feature extraction to identify key areas in the medical image to avoid misdeleting medical structure information;

[0022] The text de-identification processing includes clinical data de-identification, note de-identification, and genetic information de-identification;

[0023] The clinical data de-identification identifies patient information through a data dictionary and regular expressions, and uses a virtual ID to replace the real identifier, retaining the clinical data relationship but removing the identity mapping;

[0024] The note de-identification automatically identifies sensitive words in the note by using the natural language processing NLP method, where the sensitive information includes personal names, place names, and hospital names, replaces the identified sensitive words with common tags, and uses a semantic embedding model to maintain the semantic integrity of the original text;

[0025] The genetic information de-identification processes the patient identifier by using homomorphic encryption and a hash algorithm to anonymize the genetic information, and uses the HomomorphicEncryption technology to convert the genetic information into a state where data can still be calculated and analyzed in an encrypted state.

[0026] Preferably, the shared parameter extraction unit performs feature extraction based on the unlabeled data items, and then performs dimensionless processing to obtain a shared data set;

[0027] The shared data set includes a sensitive information identification function S, a protection degree measure H after data de-identification processing, a patient authorization identification function A, a privacy compliance identification function L, a compliance audit score F, and a valid information identification function V;

[0028] The sensitive information identification function S identifies the sensitive information existing in the unlabeled data items by using a target detection algorithm. If the sensitive information identification function S = 1, it indicates the existence; if the sensitive information identification function S = 0, it indicates the non - existence.

[0029] The protection degree measure H is obtained by calculating the ratio of the removed sensitive part to the remaining part.

[0030] The patient authorization identification function A records and queries the authorization situation of the patient for the task type T of the current i - th use through the authorization system. i If the patient authorization identification function A = 1, it indicates authorization; if the patient authorization identification function A = 0, it indicates non - authorization.

[0031] Among them, the task type T of the i - th use i represents the use for sharing.

[0032] The privacy compliance identification function L checks the compliance of the fields in the current unlabeled data item by establishing a field compliance rule base in accordance with GDPR / HIPAA requirements. If the privacy compliance identification function L = 1, it indicates compliance; if the privacy compliance identification function L = 0, it indicates non - compliance.

[0033] The compliance audit score F is obtained by scoring the unlabeled data item through a scoring system.

[0034] The valid information identification function V analyzes the remaining valid information in the unlabeled data item in reverse. If the valid information identification function V = 1, it indicates inclusion; if the valid information identification function V = 0, it indicates non - inclusion.

[0035] The encryption storage unit encrypts the unlabeled data item and the shared data set by using the AES - 256 encryption algorithm, constructs an integration of a cloud database and a data sharing platform, and stores the encrypted unlabeled data item and the shared data set in the cloud database.

[0036] Preferably, the privacy protection module includes a privacy protection analysis unit and a privacy protection evaluation unit.

[0037] The privacy protection analysis unit constructs a privacy protection degree algorithm formula, extracts the sensitive information identification function S and the protection degree measure H after data de - identification processing from the shared data set, inputs them into the privacy protection degree algorithm formula, calculates and outputs the privacy protection degree Ppr to measure the privacy protection level of the unlabeled data item.

[0038] The privacy protection degree Ppr is calculated and output through the following privacy protection degree algorithm formula;

[0039] ;

[0040] In the formula, n represents the total number of data items in the unlabeled data items, Ⅱ represents the knowledge function, e represents the exponential function, a1 represents the privacy protection change control coefficient, and S i represents the sensitive information identification function of the i-th data item in the unlabeled data items, and H i represents the protection degree measurement of the i-th data item in the unlabeled data items.

[0041] Preferably, the privacy protection evaluation unit sets the privacy protection threshold P according to the mean value specified by the user in accordance with GDPR / HIPAA regulations, and then makes a preliminary comparison and evaluation between the privacy protection threshold P and the obtained privacy protection degree Ppr to judge the privacy protection situation of the unlabeled data items requested by the current user to be shared. Based on the preliminary comparison and evaluation results, a compliance review is triggered, and the specific evaluation content is as follows;

[0042] When the privacy protection degree Ppr ≥ the privacy protection threshold P, it means that the privacy protection meets the standard, allowing entry into the data sharing process and recording it in the sharing log;

[0043] When the privacy protection degree Ppr < the privacy protection threshold P, it means that the privacy protection does not meet the standard, and at this time, a compliance review is triggered.

[0044] Preferably, the compliance review module includes a compliance analysis unit and a data validity analysis unit;

[0045] The compliance analysis unit constructs a compliance algorithm formula, extracts the patient authorization identification function A, privacy compliance identification function L, and compliance review score F of each unlabeled data item in the shared data set, inputs them into the compliance algorithm formula, calculates and outputs the compliance score Cco, and analyzes the compliance in the unlabeled data items after de-identification;

[0046] The compliance score Cco is calculated and output through the following compliance algorithm formula;

[0047] ;

[0048] In the formula, A i represents the patient authorization identification function of the i-th data item in the unlabeled data items, L i represents the privacy compliance identification function of the i-th data item in the unlabeled data items, F i is the compliance review score of the i-th data item in the unlabeled data items, e represents the exponential function, represents the compliance check intensity control coefficient, and the value is dimensionless.

[0049] Preferably, the data validity analysis unit constructs a data quality algorithm formula, extracts the valid information identification function V, inputs it into the data quality algorithm formula, calculates and outputs the data validity score Eva, and analyzes the validity in the unlabeled data items after de-identification;

[0050] The data validity score Eva is calculated and output through the following data quality algorithm formula;

[0051] ;

[0052] In the formula, V i represents the valid information identification function of the i-th data item in the unlabeled data item, represents the influence coefficient of the validity intensity of the de-identified data item.

[0053] Preferably, the data sharing real-time monitoring module includes a comprehensive analysis unit and a sharing anomaly assessment unit;

[0054] The comprehensive analysis unit comprehensively calculates the obtained compliance score Cco and data validity score Eva to obtain the data sharing index Esha, and comprehensively analyzes the compliance and validity of the shared unlabeled data items;

[0055] The data sharing index Esha is calculated and output through the following algorithm formula;

[0056] ;

[0057] In the formula, w1 and w2 respectively represent the preset weight values of the compliance score Cco and the data validity score Eva, and w1 + w2 = 1, and their specific values are set by the user.

[0058] Preferably, the sharing anomaly assessment unit sets the sharing threshold Tth by the user, evaluates the obtained data sharing index Esha with the sharing threshold Tth, analyzes the specific situation where the privacy protection of the current unlabeled data item does not meet the standard, and generates corresponding adjustments based on the specific situation. The specific assessment content is as follows;

[0059] When the data sharing index Esha ≥ the sharing threshold Tth, it indicates that the data item validity of the shared unlabeled data item is abnormal. At this time, a recovery prompt is generated to prompt the user to adaptively review the unlabeled data item and perform validity custom recovery;

[0060] When the data sharing index Esha < the sharing threshold Tth, it indicates that the data item compliance of the shared unlabeled data item is abnormal. At this time, the field compliance rule library established according to the GDPR / HIPAA requirements is rechecked, and the non-compliant fields are de-identified again.

[0061] The present invention provides an ophthalmic clinical medical data management and sharing system, which has the following beneficial effects:

[0062] (1) Through modular design, the system effectively solves many deficiencies of traditional medical data management systems in aspects such as privacy protection, data de-identification, and compliance review. The system includes a data extraction module, a de-identification module, a privacy protection module, a compliance review module, and a real-time data sharing monitoring module. Each module complements each other to ensure that privacy protection and compliance control can be achieved throughout the process from data collection, de-identification to sharing. In the data extraction module, the system can automatically extract data from the ophthalmic clinical database, perform preprocessing and standardization, enabling the subsequent de-identification and data sharing processes to proceed smoothly. In addition, the de-identification module further ensures the privacy protection of patients' sensitive information through intelligent de-identification processing and encrypted storage technology, making the shared data meet the privacy protection requirements while retaining sufficient valid information for clinical use.

[0063] (2) The privacy protection module of the system can accurately calculate and output the privacy protection degree Ppr by constructing a privacy protection degree algorithm formula, and compare it with a preset privacy protection threshold for preliminary evaluation. The working principle of this module can automatically detect whether the data meets the privacy protection standard, thereby triggering the compliance review module to ensure data compliance. The optimization effect of this design is that the system can automatically evaluate the privacy protection level before data sharing to prevent potential privacy leakage problems. The compliance review module calculates the compliance score Cco and the data validity score Eva of the de-identified data respectively through the compliance algorithm formula and the data quality algorithm formula, ensuring that the shared data not only complies with sharing rules such as GDPR and HIPAA requirements, but also has sufficient medical effectiveness. This intelligent compliance review and evaluation mechanism reduces manual intervention, improves the efficiency of the data sharing process, and at the same time reduces the risks of data leakage and non-compliance.

[0064] (3) The data sharing real-time monitoring module of this system realizes the dynamic monitoring and real-time evaluation of shared data by combining the compliance score Cco and the data validity score Eva. This module can automatically calculate the data sharing index Esha and evaluate the privacy protection and validity of shared data by setting the sharing threshold Tth. After optimization, the system can judge the abnormal situations in the data de-identification process according to the real-time sharing index. For example, when the data validity or compliance does not meet the standards, the system will automatically generate a recovery prompt, prompting the user to perform custom recovery of data validity or re-check compliance. This optimization not only ensures the efficient use of shared data, but also greatly improves the transparency and security in the data sharing process. In addition, the shared data set is encrypted and stored through the AES-256 encryption algorithm, further ensuring the security of the data and avoiding the leakage of sensitive information that may occur in the data sharing process. These optimization measures improve the security and practicality of data sharing, while ensuring the effectiveness of data in multiple scenarios such as clinical research, academic sharing, and telemedicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a schematic flowchart of an ophthalmic clinical medical data management and sharing system of the present invention;

[0066] Figure 2 It is a schematic evaluation flowchart of an ophthalmic clinical medical data management and sharing system of the present invention;

[0067] Figure 3 It is a privacy protection degree curve graph of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] Embodiment 1

[0070] Please refer to Figure 1 、 Figure 2 and Figure 3 , the present invention provides an ophthalmic clinical medical data management and sharing system. To achieve the above objectives, the present invention is realized through the following technical solutions: including a data extraction module, a de-identification module, a privacy protection module, a compliance review module, and a data sharing real-time monitoring module;

[0071] The data extraction module extracts ophthalmic clinical data by building a data sharing platform and accessing an ophthalmic clinical database, and preprocesses it to obtain standardized ophthalmic clinical data;

[0072] The de-identification module de-identifies the standardized ophthalmic clinical data by using intelligent de-identification processing, obtains un-identified data items, extracts a shared data set based on the un-identified data items, and stores it using an encryption algorithm;

[0073] The privacy protection module calculates and outputs the privacy protection degree Ppr by building a privacy protection degree algorithm formula, sets a privacy protection threshold P and makes a preliminary comparison and evaluation with the privacy protection degree Ppr, and triggers a compliance review based on the preliminary comparison and evaluation;

[0074] The compliance review module calculates and outputs a compliance score Cco and a data validity score Eva by building a compliance algorithm formula and a data quality algorithm formula, and analyzes the compliance and validity in the un-identified data items after de-identification respectively;

[0075] The data sharing real-time monitoring module calculates and outputs a data sharing index Esha by using the compliance score Cco and the data validity score Eva, sets a sharing threshold Tth and makes a sharing evaluation with the data sharing index Esha to judge the current abnormal situation of data de-identification.

[0076] In this embodiment, the system constructs a data sharing platform and accesses the ophthalmology clinical database, automatically extracts data using the API interface, and standardizes the image and text data. The data processing unit standardizes fundus images using computer vision technology, removes watermarks and patient IDs, and at the same time cleans and standardizes medical record and note data through natural language processing NLP technology to ensure the efficiency and consistency of data sharing. The de-identification module uses intelligent de-identification technology to accurately remove sensitive information such as patient identity, hospital identification, and genetic information. The image data locates the privacy area through the object detection algorithm and repairs it, retaining key medical features such as the optic disc and macula; the text data replaces the patient identifier using regular expressions and data dictionaries, and at the same time applies NLP technology to process sensitive words to ensure semantic integrity. For genetic information, the system uses homomorphic encryption and hash algorithms for encryption to prevent the patient identity from being reverse-inferred during the analysis process. The privacy protection module calculates Ppr using the privacy protection degree algorithm and compares and evaluates it with the set privacy protection threshold PPP to ensure that the data meets the privacy protection standard. Through the exponential function and the Sigmoid function, the privacy protection intensity after de-identification is evaluated to ensure compliance with privacy protection sharing rules such as GDPR and HIPAA. The compliance review module automatically analyzes the compliance and effectiveness of the de-identified data through the compliance algorithm and the data quality algorithm to ensure that the data not only meets the sharing rule requirements but also supports clinical decision-making and research applications, reducing the loopholes and delays in manual review. The data sharing real-time monitoring module generates the data sharing index Esha through the compliance score and the data effectiveness score, and compares it with the preset sharing threshold Tth in real time to monitor abnormal situations during the data sharing process. If an abnormality is found, the system will automatically prompt the user to perform data recovery or re-evaluation to ensure the security and compliance of data sharing. Compared with traditional technologies, the system significantly improves the security, compliance, and effectiveness of ophthalmology data sharing through intelligent de-identification, automated compliance review, real-time monitoring mechanism, and encrypted storage, ensuring the balance between privacy protection and data integrity, and promoting the wide application of medical data sharing in academic research, telemedicine, and clinical decision-making.

[0077] Embodiment 2

[0078] Please refer to Figure 1 , specifically: The data extraction module includes a data extraction unit and a data processing unit;

[0079] The data extraction unit constructs a data sharing platform, sets the API application program interface, accesses the ophthalmology clinical database, and extracts the ophthalmology clinical data sent by the user in the form of a sharing request after the user sends the sharing request, and transmits it to the data sharing platform;

[0080] The ophthalmic clinical data includes image data and text data, where the text data includes clinical data, note data, and genetic information;

[0081] The data processing unit preprocesses the extracted ophthalmic clinical data, and the preprocessing includes image data processing, text data processing, data cleaning, and standardization to obtain standardized ophthalmic clinical data;

[0082] The image data processing performs standardization processing on fundus images and retinal scans using computer vision technology to remove redundant information such as watermarks and patient IDs;

[0083] The text data processing processes notes and medical records using natural language processing NLP technology to identify sensitive information such as names and diagnoses;

[0084] The data cleaning and standardization denoise and standardize all ophthalmic clinical data and unify the data format.

[0085] In this embodiment, the data extraction module of the system includes a data extraction unit and a data processing unit, which improves the sharing efficiency and security of ophthalmic clinical data through a refined processing process. The data extraction unit constructs a data sharing platform and sets up an API interface, which can efficiently access the ophthalmic clinical database, extract the image data and text data requested by the user, and ensure the automation and real-time nature of the data extraction process. The image data includes fundus images and retinal scans, and the text data covers clinical data, medical record notes, and genetic information. The implementation of this module ensures the unified extraction and rapid processing of different types of data. Then, the data processing unit comprehensively preprocesses the extracted ophthalmic clinical data, including image data processing, text data processing, and data cleaning and standardization. The image data is standardized through computer vision technology to remove redundant information such as watermarks and patient IDs, and at the same time, natural language processing NLP technology is applied to identify and remove sensitive information from the text data to ensure de-identification without losing medical value. In addition, data cleaning and standardization technologies make all data reach a unified specification through operations such as denoising and format unification, providing standardized data support for subsequent sharing and analysis.

[0086] Embodiment 3

[0087] Please refer to Figure 1 , specifically: the de-identification module includes an intelligent de-identification processing unit, a shared parameter extraction unit, and an encrypted storage unit;

[0088] The intelligent de-identification processing unit performs de-identification processing on the standardized ophthalmic clinical data by setting intelligent de-identification, removes sensitive information from the standardized ophthalmic clinical data, and obtains unlabeled data items. The intelligent de-identification processing includes image data de-identification processing and text de-identification processing;

[0089] The sensitive information includes privacy areas, patient information, sensitive words, and genetic information;

[0090] The image data de-identification processing uses a target detection algorithm to locate privacy areas in the image, such as name watermarks and hospital information in the corners of the image, and then uses an image repair algorithm to perform seamless replacement of the privacy areas. At the same time, edge detection Canny, SIFT, and SURF feature extraction are used to identify key areas in the medical image to avoid accidental deletion of medical structure information;

[0091] Technical effect: Accurately remove directly exposed identity information in the image and retain medical feature areas in the ophthalmic image, such as the optic disc and macula, to ensure that medical analysis is not affected;

[0092] The text de-identification processing includes clinical data de-identification, note de-identification, and genetic information de-identification;

[0093] The clinical data de-identification identifies patient information, such as name, ID card, mobile phone, and address fields, through a data dictionary and regular expressions, and then uses a virtual ID, such as U123456, to replace the real identifier, retaining the clinical data relationship but removing the identity mapping;

[0094] Technical effect: Protect patient identity information from being reverse-inferred, maintaining data integrity and availability, such as data can be traced by patient ID but the identity cannot be reverse-checked;

[0095] The note de-identification automatically identifies sensitive words in the note by using natural language processing (NLP) methods. The sensitive information includes personal names, place names, and hospital names, and replaces the identified sensitive words with common tags, and uses a semantic embedding model to maintain the semantic integrity of the original text;

[0096] Technical effect: Automatically identify privacy fields in doctors' natural language writing and protect text-based privacy information to the greatest extent without affecting semantic integrity;

[0097] The genetic information de-identification processes the patient identifier by using homomorphic encryption and hash algorithms to anonymize the genetic information. Using Homomorphic Encryption technology, the genetic information is converted into a state where data can still be calculated and analyzed under encryption, but the individual identity cannot be restored. The family history information is processed by structural perturbation to blur it, avoiding inferring individual identity from the family structure;

[0098] Technical function: Prevent highly sensitive genetic data from being directly associated with patient identities, and support data analysis and research in an encrypted state, such as mutation frequency analysis and genetic pathway calculation.

[0099] The shared parameter extraction unit extracts features based on unlabeled data items, performs dimensionless processing, and obtains a shared data set.

[0100] The shared data set includes a sensitive information identification function S, a protection degree measure H after data de-identification processing, a patient authorization identification function A, a privacy compliance identification function L, a compliance audit score F, and a valid information identification function V.

[0101] The sensitive information identification function S uses a target detection algorithm to identify sensitive information existing in unlabeled data items. If the sensitive information identification function S = 1, it indicates the existence; if the sensitive information identification function S = 0, it indicates the non-existence.

[0102] The protection degree measure H is obtained by calculating the ratio of the removed sensitive part to the remaining part. For example, for image data, it is based on the ratio of the removed sensitive area to the total image area, and for text data, it is based on the number of processed fields and the total number of fields containing sensitive information.

[0103] The patient authorization identification function A records the authorization situation of the query patient for the task type T of the current i-th use through an authorization system. i If the patient authorization identification function A = 1, it indicates authorization; if the patient authorization identification function A = 0, it indicates non-authorization.

[0104] Among them, the task type T of the i-th use i represents the use for sharing, such as scientific research use, remote diagnosis and treatment use.

[0105] The privacy compliance identification function L checks the compliance of fields in the current unlabeled data items by establishing a field compliance rule library in accordance with GDPR / HIPAA requirements. If the privacy compliance identification function L = 1, it indicates compliance; if the privacy compliance identification function L = 0, it indicates non-compliance.

[0106] The compliance audit score F is obtained by scoring the unlabeled data items through a scoring system.

[0107] The valid information identification function V reversely analyzes the remaining valid information in the unlabeled data items. If the valid information identification function V = 1, it indicates inclusion, that is, the data contains medically or research-useful content, such as retinal images, diagnostic fields, and treatment labels, etc. If the valid information identification function V = 0, it indicates non-inclusion, that is, the data has lost effectiveness due to de-identification or other processing.

[0108] The encrypted storage unit encrypts the unlabeled data items and the shared data set by using the AES-256 encryption algorithm, constructs an integration of the cloud database and the data sharing platform, and stores the encrypted unlabeled data items and the shared data set in the cloud database.

[0109] In this embodiment, the de-identification module of the system consists of an intelligent de-identification processing unit, a shared parameter extraction unit, and an encrypted storage unit, and a complete and efficient implementation mechanism is carried out around data privacy protection, information validity retention, and compliant sharing. The intelligent de-identification processing unit performs de-identification operations on the classified image and text data for the standardized ophthalmic clinical data. In terms of image data, the system accurately identifies the privacy areas in the image through the object detection algorithm, and uses the image repair technology for seamless replacement. At the same time, combined with edge detection and feature extraction algorithms such as Canny, SIFT, and SURF, it ensures that key medical areas such as the optic disc and macula are not accidentally deleted, so as to retain the clinical diagnostic value while removing sensitive information. In terms of text data, the system identifies and replaces sensitive content such as names, ID numbers, and addresses in the clinical fields through the data dictionary and regular expressions, and uses NLP and semantic embedding models to process sensitive words such as people's names and place names in the notes to ensure semantic integrity. For genetic information, the system introduces homomorphic encryption and hash algorithms to encrypt and store sensitive gene information, and uses the structure perturbation method to prevent identity reverse inference to achieve research and analysis in the encrypted state. The shared parameter extraction unit extracts key indicators after the data de-identification, performs dimensionless processing on them, and generates a standardized shared data set. These indicators not only provide a quantitative basis for subsequent privacy protection and compliance assessment, but also ensure the flexible application of data in multiple tasks and scenarios. The encrypted storage unit uses the AES-256 encryption algorithm to securely encrypt the processed unlabeled data and the shared data set, and integrates the cloud database and the sharing platform to achieve full-process encrypted storage and anti-leakage management of the data before and after sharing.

[0110] Edge detection processing flow in image data processing:

[0111] Image preprocessing: First, convert the input ophthalmic image, such as a fundus image, to grayscale, and use the Canny edge detection algorithm in OpenCV to extract the structural edge information of the image. The parameter thresholds used are 50 / 150, which is a common and effective configuration method of this algorithm in medical image processing;

[0112] Contour Recognition and Region Judgment: By performing dilation processing and contour extraction on the edge results, possible sensitive regions in the image are recognized, such as watermarks in the lower right corner, corner labels, etc. Empirical judgment rules are set. For example, when the width of the recognized region is greater than 80, the height is less than 100, and the region is located on the right or lower side of the image, with x > 500 and y > 500, it is determined as a privacy region to be removed;

[0113] Key Region Protection Mechanism: The purpose of this edge detection mechanism is to accurately distinguish between privacy regions and medical structure regions, avoiding direct covering or blurring processing from affecting the image quality;

[0114] Image Restoration and Seamless Replacement: After identifying the privacy region, the cv2.inpaint() restoration algorithm is used for seamless filling. The method used is INPAINT_TELEA. This algorithm supports natural edge extension, can preserve the consistency of image texture, and enhance the medical readability of the image.

[0115] NLP Sensitive Word Recognition and Desensitization Process in Text Data Processing:

[0116] For possible patient privacy information in doctor's notes and clinical description texts, the system of this application adopts a sensitive word recognition and semantic protection replacement mechanism based on the natural language processing NLP method. Its processing steps include:

[0117] Entity Recognition Strategy: Use a rule model based on regular expressions and a named template recognition method to perform high-precision matching on personal names, hospital names, and place names in Chinese texts; at the same time, recognize structured fields, such as 18-digit ID numbers and 11-digit mobile phone numbers, and use regular patterns \d{18} and 1[3-9]\d{9} for desensitization replacement;

[0118] Tagging Processing Mechanism: Replace the recognized sensitive entities with standard tags, such as <PER_MASK>, <LOC_MASK>, <ORG_MASK>, etc., while retaining the original grammatical structure of the sentence;

[0119] Semantic Consistency Assurance: By restricting the word length, position, and context window length after entity replacement, the preservation of semantic coherence is achieved, ensuring that the text can still be used clinically or for algorithm analysis after desensitization processing.

[0120] Embodiment 4

[0121] Please refer to Figure 1 、 Figure 2 and Figure 3 Specifically: The privacy protection module includes a privacy protection analysis unit and a privacy protection evaluation unit;

[0122] The privacy protection analysis unit constructs a privacy protection degree algorithm formula, extracts the sensitive information identification function S and the protection degree metric H after data de-identification in the shared dataset, inputs them into the privacy protection degree algorithm formula, calculates and outputs the privacy protection degree Ppr to measure the privacy protection level of the un-identified data items;

[0123] The privacy protection degree Ppr is calculated and output through the following privacy protection degree algorithm formula;

[0124] ;

[0125] In the formula, n represents the total number of data items in the un-identified data items, Ⅱ represents the knowledge function, also known as the 0-1 function, which is used to map whether a certain condition holds to 0 or 1, e represents the exponential function, a1 represents the privacy protection change control coefficient, S i represents the sensitive information identification function of the i-th data item in the un-identified data items, H i represents the protection degree metric of the i-th data item in the un-identified data items, represents the Sigmoid function, which controls how the degree of de-identification is converted into the protection degree;

[0126] Specific example, assume there are 5 data items in the un-identified data items, namely medical images, doctor's notes, test forms, no sensitive data, and medical images. The sensitive information identification functions S are respectively: 1, 1, 1, 1, 0, 1, and the protection degree metrics H after data de-identification are respectively: 0.8, 0.6, 0.3, 0.0, 0.95. Set a1 = 5 and substitute it into the formula for calculation and output;

[0127] Ppr = 1 / 5(0.180 + 0.0474 + 0.1824 + 0 + 0.0085) = 0.0513.

[0128] The privacy protection evaluation unit allows the user to set the privacy protection threshold P according to the mean value specified by GDPR / HIPAA regulations. For example, the allowable value of GDPR is ≥0.75, and the running value of HIPAA is ≥0.8. Therefore, the privacy protection threshold P = (0.7 + 0.8) / 2 = 0.77, taking two decimal places. Then, the privacy protection threshold P is compared and evaluated with the obtained privacy protection degree Ppr to judge the privacy protection situation of the un-identified data items requested by the current user for sharing, and based on the preliminary comparison and evaluation results, trigger a compliance review. The specific evaluation content is as follows;

[0129] When the privacy protection degree Ppr ≥ the privacy protection threshold P, it means that the privacy protection meets the standard, allowing entry into the data sharing process and recording it in the sharing log;

[0130] When the privacy protection degree Ppr < the privacy protection threshold P, it indicates that the privacy protection fails to meet the standard, and at this time, a compliance review is triggered.

[0131] In this embodiment, the privacy protection module of the system consists of a privacy protection analysis unit and a privacy protection evaluation unit, aiming to ensure that ophthalmic data strictly complies with privacy protection standards during the sharing process and safeguard patient privacy from being leaked through an intelligent evaluation and review mechanism. The privacy protection analysis unit calculates the privacy protection degree Ppr of the unlabeled data items by constructing a privacy protection degree algorithm. This algorithm is based on the sensitive information identification function S and the degree of protection for data de-identification measurement H, comprehensively considering the sensitivity and de-identification degree of the data items. Through the combination of the Sigmoid function and the exponential function, the system dynamically evaluates the privacy protection intensity of each data item, thereby outputting a comprehensive privacy protection degree to reflect the privacy protection level of the data. The privacy protection evaluation unit sets the privacy protection threshold P according to the privacy protection standards stipulated by GDPR and HIPAA. For example, the average value of the two is taken as the privacy protection threshold P. When the calculated privacy protection degree Ppr is greater than or equal to the threshold P, the system considers that the data privacy protection meets the standard, allows it to enter the data sharing process, and records it in the sharing log; conversely, when Ppr is less than the threshold, the system will trigger a compliance review to further verify the compliance of the data and privacy protection measures, ensuring the legality and security of data sharing. The implementation of this module not only improves the accuracy and efficiency of the privacy protection process through automated evaluation and intelligent analysis but also reduces the risk of manual intervention by comparing with the privacy protection threshold.

[0132] Embodiment 5

[0133] Please refer to Figure 1 , specifically: The compliance review module includes a compliance analysis unit and a data validity analysis unit;

[0134] The compliance analysis unit calculates the compliance score Cco by constructing a compliance algorithm formula, extracting the patient authorization identification function A, privacy compliance identification function L, and compliance review score F of each unlabeled data item in the shared dataset, and inputting them into the compliance algorithm formula for analysis of the compliance in the de-identified unlabeled data items;

[0135] The compliance score Cco is calculated and output through the following compliance algorithm formula;

[0136] ;

[0137] where A i represents the patient authorization identification function of the i-th data item in the unlabeled data item, L i represents the privacy compliance identification function of the i-th data item in the unlabeled data item, F iThe compliance audit score of the i-th data item in the unlabeled data item, where e represents the exponential function, represents the compliance check intensity control coefficient, which is dimensionless, and the larger the value, the stricter the compliance check.

[0138] The data validity analysis unit constructs a data quality algorithm formula, extracts the valid information identification function V, inputs it into the data quality algorithm formula, calculates and outputs the data validity score Eva, and analyzes the validity in the unlabeled data item after de-identification;

[0139] The data validity score Eva is calculated and output through the following data quality algorithm formula;

[0140] ;

[0141] In the formula, V i represents the valid information identification function of the i-th data item in the unlabeled data item, represents the influence coefficient of the validity strength of the data item after de-identification. The larger the value, the greater the loss of validity after de-identification.

[0142] In this embodiment, the compliance analysis unit of the system constructs a compliance algorithm, extracts the patient authorization identifier A, privacy compliance identifier L, and compliance audit score F of each unlabeled data item, inputs them into the compliance algorithm formula, calculates the compliance score Cco, and evaluates the compliance of the data. This score combines patient authorization, privacy compliance, and compliance audit scores to ensure that all shared data is processed in compliance within the framework of sharing rules. By controlling the compliance check intensity through the exponential function, the system can provide a more strict or lenient review based on the compliance score, flexibly responding to different types of data sharing requirements. At the same time, the data validity analysis unit constructs a data quality algorithm, extracts the valid information identification function V of each unlabeled data item, and calculates the data validity score Eva to evaluate the validity of the data item after de-identification. This analysis considers the degree of validity loss during the data de-identification process to ensure that even after de-identification, the data still retains sufficient medical and research value. This scoring mechanism effectively measures the balance between data validity and the de-identification process, ensuring that while protecting privacy, the data can still support clinical decision-making and academic research. Through dual analysis of compliance and validity, this module realizes the comprehensive evaluation of compliance and validity during the data sharing process and automatically provides review support for data sharing.

[0143] Example 6

[0144] Please refer to Figure 1 and Figure 2 Specifically: The data sharing real-time monitoring module includes a comprehensive analysis unit and a sharing exception evaluation unit;

[0145] The comprehensive analysis unit comprehensively calculates the data sharing indicator Esha by using the obtained compliance score Cco and data validity score Eva, and comprehensively analyzes the compliance and validity of the unlabeled data items for sharing;

[0146] The data sharing indicator Esha is calculated and output through the following algorithm formula;

[0147] ;

[0148] In the formula, w1 and w2 respectively represent the preset weight values of the compliance score Cco and the data validity score Eva, and w1 + w2 = 1. Their specific values are set by the user.

[0149] The sharing anomaly assessment unit sets the sharing threshold Tth by the user, and conducts a sharing assessment by comparing the obtained data sharing indicator Esha with the sharing threshold Tth, analyzes the specific situation where the privacy protection of the current unlabeled data item does not meet the standard, and generates corresponding adjustments based on the specific situation. The specific assessment content is as follows;

[0150] When the data sharing indicator Esha ≥ the sharing threshold Tth, it indicates that the data item validity of the unlabeled data item for sharing is abnormal. At this time, a recovery prompt is generated to prompt the user to adaptively review the unlabeled data item and perform validity custom recovery;

[0151] When the data sharing indicator Esha < the sharing threshold Tth, it indicates that the data item compliance of the unlabeled data item for sharing is abnormal. At this time, the field compliance rule library established according to the GDPR / HIPAA requirements is re - used to conduct a compliance re - check, and secondary de - identification is performed on the non - compliant fields.

[0152] In this embodiment, the real-time monitoring module for data sharing of the system includes a comprehensive analysis unit and a sharing anomaly evaluation unit, aiming to ensure compliance and effectiveness during the data sharing process and promptly detect abnormal situations for automatic adjustment. The comprehensive analysis unit generates a data sharing index Esha by comprehensively calculating the compliance score Cco and the data effectiveness score Eva, and comprehensively analyzes the compliance and effectiveness of the shared data. Through this method, the system can dynamically evaluate the comprehensive quality of the shared data and provide decision-making support for subsequent data sharing processes. The sharing anomaly evaluation unit conducts real-time comparison and evaluation of the calculated sharing index according to the sharing threshold Tth set by the user. When the data sharing index Esha is greater than or equal to the sharing threshold, it indicates an abnormality in data effectiveness. The system will generate a recovery prompt to remind the user to conduct an adaptive review of the data and perform custom recovery of effectiveness. If the data sharing index is less than the threshold, it indicates an abnormality in compliance. The system will re-check the compliance of the data according to sharing rules such as GDPR and HIPAA and perform secondary de-identification processing on non-compliant fields.

[0153] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention.

Claims

1. An ophthalmology clinical medical data management and sharing system, characterized in that: It includes a data extraction module, a de-identification module, a privacy protection module, a compliance review module, and a real-time data sharing monitoring module; The data extraction module constructs a data sharing platform, accesses the ophthalmology clinical database, extracts ophthalmology clinical data, and performs preprocessing to obtain standardized ophthalmology clinical data; The de-identification module performs de-identification on the standardized ophthalmology clinical data through intelligent de-identification processing, obtains unlabeled data items, extracts a shared data set based on the unlabeled data items, and then stores it using an encryption algorithm; The privacy protection module constructs a privacy protection degree algorithm formula, calculates and outputs the privacy protection degree Ppr, sets a privacy protection threshold P and compares it with the privacy protection degree Ppr for a preliminary evaluation, and triggers a compliance review based on the preliminary evaluation; The compliance review module constructs a compliance algorithm formula and a data quality algorithm formula, calculates and outputs a compliance score Cco and a data validity score Eva, and analyzes the compliance and validity in the de-identified unlabeled data items respectively; The real-time data sharing monitoring module calculates and outputs a data sharing index Esha through the compliance score Cco and the data validity score Eva, sets a sharing threshold Tth and conducts a sharing evaluation with the data sharing index Esha to judge the current abnormal situation of data de-identification; 2. An ophthalmology clinical medical data management and sharing system according to claim 1, characterized in that: The data extraction module includes a data extraction unit and a data processing unit; The data extraction unit constructs a data sharing platform, sets an API application program interface, accesses the ophthalmology clinical database, and after the user sends a sharing request, extracts the ophthalmology clinical data sent by the user and transmits it to the data sharing platform; The ophthalmology clinical data includes image data and text data, where the text data includes clinical data, note data, and genetic information; The data processing unit performs preprocessing on the extracted ophthalmology clinical data, and the preprocessing includes image data processing, text data processing, data cleaning, and standardization, to obtain standardized ophthalmology clinical data; The image data processing standardizes fundus images and retinal scan images using computer vision technology to remove redundant information; The text data processing processes notes and medical records using natural language processing NLP technology to identify sensitive information therein; The data cleaning and standardization denoise and standardize all ophthalmology clinical data and unify the data format; 3. An ophthalmic clinical medical data management and sharing system according to claim 2, characterized in that: The de-identification module includes an intelligent de-identification processing unit, a shared parameter extraction unit, and an encrypted storage unit; The intelligent de-identification processing unit performs de-identification processing on the standardized ophthalmology clinical data through setting intelligent de-identification processing, removes sensitive information in the standardized ophthalmology clinical data, and obtains unlabeled data items. The intelligent de-identification processing includes image data de-identification processing and text de-identification processing; The sensitive information includes privacy areas, patient information, sensitive words, and genetic information; The de-identification process of the image data uses a target detection algorithm to locate the privacy areas in the image, and then uses an image inpainting algorithm to seamlessly replace the privacy areas. At the same time, edge detection Canny, SIFT, and SURF feature extraction are used to identify the key areas in the medical image to avoid accidentally deleting medical structure information; The de-identification process of the text includes the de-identification of clinical data, the de-identification of notes, and the de-identification of genetic information; The de-identification of the clinical data uses a data dictionary and regular expressions to identify patient information, and then uses a virtual ID to replace the real identifier, retaining the clinical data relationship but removing the identity mapping; The de-identification of the notes uses natural language processing (NLP) methods to automatically identify sensitive words in the notes. The sensitive information includes personal names, place names, and hospital names, and the identified sensitive words are replaced with common tags, and a semantic embedding model is used to maintain the semantic integrity of the original text; The de-identification of the genetic information processes the patient identifier using homomorphic encryption and hash algorithms to anonymize the genetic information. Using the Homomorphic Encryption technology, the genetic information is transformed into a state where data can still be computationally analyzed while encrypted.

4. An ophthalmic clinical medical data management and sharing system according to claim 3, characterized in that: The shared parameter extraction unit extracts features based on the unlabeled data items, and then performs dimensionless processing to obtain a shared data set; The shared data set includes a sensitive information identification function S, a protection degree measure H after the data de-identification process, a patient authorization identification function A, a privacy compliance identification function L, a compliance audit score F, and a valid information identification function V; The sensitive information identification function S uses a target detection algorithm to identify the sensitive information existing in the unlabeled data items. If the sensitive information identification function S = 1, it means it exists; if the sensitive information identification function S = 0, it means it does not exist; The protection degree measure H is obtained by calculating the ratio of the removed sensitive part to the remaining part; The patient authorization identification function A records and queries, through the authorization system, the authorization status of the patient for the task type T of the current i-th use i If the patient authorization identification function A = 1, it indicates authorization; if the patient authorization identification function A = 0, it indicates non-authorization; Among them, the task type T of the i-th use i indicates the use for sharing; The privacy compliance identification function L checks the compliance of the fields in the current unlabeled data items by referring to the GDPR / HIPAA requirements to establish a field compliance rule base. If the privacy compliance identification function L = 1, it means it is compliant; if the privacy compliance identification function L = 0, it means it is non-compliant; The compliance audit score F is obtained by scoring the unlabeled data items through a scoring system; The valid information identification function V analyzes the remaining valid information in the unlabeled data items in reverse. If the valid information identification function V = 1, it means it contains; if the valid information identification function V = 0, it means it does not contain; The encryption storage unit uses the AES-256 encryption algorithm to encrypt the unlabeled data items and the shared data set, and constructs a cloud database and integrates it with the data sharing platform to store the encrypted unlabeled data items and the shared data set in the cloud database.

5. An ophthalmic clinical medical data management and sharing system according to claim 4, characterized in that: The privacy protection module includes a privacy protection analysis unit and a privacy protection evaluation unit; The privacy protection analysis unit constructs a privacy protection degree algorithm formula, extracts the sensitive information identification function S in the shared dataset and the protection degree measure H after data de-identification processing, inputs them into the privacy protection degree algorithm formula, calculates and outputs the privacy protection degree Ppr, and measures the privacy protection level of the non-identified data items. The privacy protection degree Ppr is calculated and output through the following privacy protection degree algorithm formula. ; Wherein, n represents the total number of data items in the unlabeled data items, Ⅱ represents the knowledge function, e represents the exponential function, a1 represents the privacy protection change control coefficient, S i represents the sensitive information identification function of the i-th data item in the unlabeled data items, H i represents the protection degree measurement of the i-th data item in the unlabeled data items, represents the Sigmoid function.

6. The ophthalmology clinical medical data management and sharing system according to claim 5, wherein: The privacy protection evaluation unit allows the user to set a privacy protection threshold P according to the mean value specified by GDPR / HIPAA, and then makes a preliminary comparison and evaluation between the privacy protection threshold P and the obtained privacy protection degree Ppr to judge the privacy protection situation of the non-identified data items requested by the current user for sharing. Based on the preliminary comparison and evaluation results, a compliance review is triggered. The specific evaluation content is as follows. When the privacy protection degree Ppr ≥ the privacy protection threshold P, it indicates that the privacy protection meets the standard, allowing entry into the data sharing process and recording it in the sharing log. When the privacy protection degree Ppr < the privacy protection threshold P, it indicates that the privacy protection does not meet the standard, and at this time, a compliance review is triggered.

7. An ophthalmic clinical medical data management and sharing system according to claim 6, characterized in that: The compliance review module includes a compliance analysis unit and a data validity analysis unit. The compliance analysis unit constructs a compliance algorithm formula, extracts the patient authorization identification function A, the privacy compliance identification function L, and the compliance review score F of each non-identified data item in the shared dataset, inputs them into the compliance algorithm formula, calculates and outputs the compliance score Cco, and analyzes the compliance in the non-identified data items after de-identification. The compliance score Cco is calculated and output through the following compliance algorithm formula. ; Where, A i represents the patient authorization identification function of the i-th data item in the unlabeled data item, L i represents the privacy compliance identification function of the i-th data item in the unlabeled data item, F i is the compliance audit score of the i-th data item in the unlabeled data item, e represents the exponential function, represents the compliance check intensity control coefficient, and the value is dimensionless.

8. An ophthalmic clinical medical data management and sharing system according to claim 7, characterized in that: The data validity analysis unit constructs a data quality algorithm formula, extracts the valid information identification function V, inputs it into the data quality algorithm formula, calculates and outputs the data validity score Eva, and analyzes the validity in the non-identified data items after de-identification. The data validity score Eva is calculated and output through the following data quality algorithm formula. ; Where, V i represents the valid information identification function of the i-th data item in the unlabeled data item, represents the influence coefficient of the validity strength of the de-identified data item.

9. An ophthalmic clinical medical data management and sharing system according to claim 1, characterized in that: The data sharing real-time monitoring module includes a comprehensive analysis unit and a sharing anomaly evaluation unit. The comprehensive analysis unit comprehensively calculates the obtained compliance score Cco and data validity score Eva to obtain the data sharing index Esha, and comprehensively analyzes the compliance and validity of sharing non-identified data items. The data sharing index Esha is calculated and output through the following algorithm formula. ; In the formula, w1 and w2 respectively represent the preset weight values of the compliance score Cco and the data validity score Eva, and w1 + w2 = 1, and their specific values are set by the user.

10. An ophthalmic clinical medical data management and sharing system according to claim 9, characterized in that: The sharing anomaly evaluation unit allows the user to set a sharing threshold Tth, makes a sharing evaluation between the obtained data sharing index Esha and the sharing threshold Tth, analyzes the specific situation where the privacy protection of the current non-identified data item does not meet the standard, and generates corresponding adjustments based on the specific situation. The specific evaluation content is as follows. When the data sharing metric Esha ≥ the sharing threshold Tth, it indicates an abnormality in the data item validity of sharing un-identified data items. At this time, a recovery prompt is generated to prompt the user to adaptively review the un-identified data items and perform custom validity recovery. When the data sharing metric Esha < the sharing threshold Tth, it indicates an abnormality in the data item compliance of sharing un-identified data items. At this time, the compliance re-check is carried out again based on the field compliance rule library established according to the GDPR / HIPAA requirements, and secondary de-identification is performed on the non-compliant fields.

Citation Information

Patent Citations

  • Fine-grained security data sharing method for patient health record privacy protection

    CN116663047A

  • Medical data security sharing method and system based on block chain

    CN119357995A