Medical record copying and borrowing authority management system and method

By dynamically setting desensitization rules for medical record copies, and generating dynamic masking rules based on the risk of re-identification of medical record data, the problem of the mismatch between privacy protection and the usability of scientific research data in traditional methods is solved, and a balance between risk control and usability is achieved.

CN121659355AInactive Publication Date: 2026-03-13重庆市渝北区人民医院
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-03-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional methods for desensitizing medical record data are insufficient to effectively address the varying risks of re-identification among different medical records, leading to a mismatch between the strength of privacy protection and the usability of research data, and posing a risk of privacy leakage.

Method used

By obtaining the correlation strength between different fields in the medical record dataset and the patient's identity, the risk contribution and sensitivity are determined, dynamic masking rules are generated, and desensitization processing is performed when copying medical records. Access authorization is then granted based on the completeness of the core content of the desensitization results.

Benefits of technology

It achieves precise risk control during medical record copying, avoiding the problems of insufficient protection for high-risk cases or excessive desensitization for low-risk cases in traditional methods, ensuring privacy and preserving research value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121659355A_ABST
    Figure CN121659355A_ABST
Patent Text Reader

Abstract

The invention provides a medical record copying and borrowing authority management system and method. The risk contribution degree of each field during patient identity re-identification is determined through the association strength between different fields in a medical record data set and corresponding patient identities; performing risk coupling based on the risk contribution degrees of all the fields during patient identity re-identification in combination with the field sensitivity to obtain re-identification risk indexes of different medical record records, and further generating a dynamic mask rule according to all the re-identification risk indexes; performing desensitization processing on the medical record data set through the dynamic mask rule, and determining the dynamic information availability rate of the medical record data set after the desensitization processing according to the integrity of core contents of different medical record records in a desensitization result; and performing access authorization on batch copying of the medical record data set according to the dynamic information availability rate in combination with a preset permission verification rule. By adopting the scheme of the invention, the desensitization rule during the copying of the medical record can be dynamically set based on the re-identification risk of the medical record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of access control technology, and more specifically, to a medical record copying and borrowing access control system and method. Background Technology

[0002] Access control refers to the process of transforming the disordered application behavior of users or entities to access system resources into ordered authorization through identity authentication mechanisms, access control policies, permission allocation rules and dynamic auditing methods. This prevents unauthorized access to resources, unauthorized operation by users, abuse or malicious tampering of sensitive information, and ultimately ensures the secure and controllable management of system resources.

[0003] Medical record copying and borrowing access management refers to a management mechanism that regulates the copying and borrowing of medical records through identity verification mechanisms, hierarchical access rules, borrowing approval processes, and usage traceability auditing. It clarifies the scope, purpose, and duration of medical record use, thereby preventing unauthorized access, excessive copying and borrowing, misuse of patient privacy data, or content tampering, and avoiding medical privacy leaks and medical record data security risks. Traditional medical record data anonymization methods often rely on fixed anonymization rules, which are insufficient to effectively address the varying risks of re-identification among different medical records. This leads to a mismatch between the strength of privacy protection and the actual needs of research data usability. For example, traditional anonymization might only retain the year of birth for rare disease cases (such as hereditary angioedema). Even so, combining the fields "hereditary angioedema + age range of onset + hospital visited" still allows for locating specific patients in a regional patient database, leading to privacy leaks due to re-identification during bulk copying. Therefore, dynamically setting anonymization rules for medical record copying based on the risk of re-identification has become a challenge for the industry. Summary of the Invention

[0004] This application provides a medical record copying and borrowing permission management system and method, which can dynamically set desensitization rules for medical record copying based on the risk of medical record re-identification.

[0005] Firstly, this application provides a method for de-identifying medical record data, used in a medical record copying and borrowing access management system to authorize copying access to de-identified medical record data. The method includes the following steps: Obtain the medical record dataset when batch copying medical records for scientific research projects. The medical record dataset includes medical records of multiple patients and multiple fields in each medical record. The risk contribution of each field in patient identity re-identification is determined by the correlation strength between different fields in the medical record dataset and the corresponding patient identities. Risk coupling is performed based on the risk contribution of all fields in patient identity re-identification and field sensitivity to obtain the re-identification risk index of different medical records. Then, dynamic masking rules for copying medical records are generated based on all re-identification risk indices. The medical record dataset is anonymized using the dynamic masking rules, and the dynamic information availability rate of the medical record dataset after anonymization is determined based on the completeness of the core content of different medical record records in the anonymization results. Access authorization is granted for batch copying of the medical record dataset based on the availability of the dynamic information and the preset permission verification rules.

[0006] In some embodiments, determining the risk contribution of each field in patient identity re-identification based on the association strength between different fields in the medical record dataset and the corresponding patient identity specifically includes: Obtain unique identifiers for different patients in the medical record dataset; The strength of the association between different fields and corresponding patient identities in the medical record dataset is determined by all unique identifiers; Determine the re-identification risk characteristics of different fields in the medical record dataset; The risk contribution of each field in patient identity re-identification is determined based on all re-identification risk characteristics and all association strengths.

[0007] In some embodiments, risk coupling is performed based on the risk contribution of all fields in patient identity re-identification, combined with field sensitivity, to obtain a re-identification risk index for different medical records. Specifically, this includes: Determine the field sensitivity of different fields in the medical record dataset; Select one medical record as the selected medical record; Retrieve all fields corresponding to the selected medical record; The risk contribution of all fields in patient identity re-identification and the field sensitivity of all fields are coupled into risk factors for each field; The re-identification risk index of selected medical records is determined based on all risk factors; Continue to determine the re-identification risk index for the remaining medical records.

[0008] In some embodiments, the dynamic masking rules for generating photocopied medical records based on all re-identification risk indices specifically include: Set dynamic risk thresholds when photocopying medical records; The dynamic risk threshold and all re-identification risk indices are used to divide all medical records into sets of medical records with different risk levels. Dynamic masking rules for generating photocopied medical records based on sets of medical records with different risk levels.

[0009] In some embodiments, the desensitization process of the medical record dataset using the dynamic masking rules specifically includes: Obtain the risk level corresponding to different medical records in the medical record dataset; An adaptive masking scheme is generated for each medical record based on the dynamic masking rules and the risk levels corresponding to different medical records. The medical record dataset is anonymized based on an adaptive masking scheme for each medical record.

[0010] In some embodiments, determining the dynamic information availability rate of the medical record dataset after desensitization based on the completeness of the core content of different medical record records in the desensitization results specifically includes: Obtain a list of core information corresponding to the research project; The completeness of the core content of different medical record entries in the desensitization results is determined by the core information list; The availability of the anonymized medical record dataset is verified based on the completeness of the core content of different medical record records, thus obtaining the dynamic information availability rate of the anonymized medical record dataset.

[0011] In some embodiments, authorizing access to the batch copying of the medical record dataset based on the dynamic information availability rate and preset permission verification rules specifically includes: The desensitization process of the medical record dataset is dynamically adjusted based on the availability of the dynamic information. Based on preset permission verification rules and adjustment results, a permission certificate for batch copying of the medical record dataset is generated; Access is granted for batch copying of the medical record dataset based on the aforementioned permission credentials.

[0012] Secondly, this application provides a medical record copying and borrowing access management system, including a medical record data desensitization unit, wherein the medical record data desensitization unit includes: The acquisition module is used to acquire the medical record dataset when batch copying medical records for scientific research projects. The medical record dataset includes medical records of multiple patients and multiple fields in each medical record. The processing module is used to determine the risk contribution of each field in the patient identity re-identification process by the correlation strength between different fields in the medical record dataset and the corresponding patient identity. The processing module is also used to perform risk coupling based on the risk contribution of all fields in patient identity re-identification and the field sensitivity to obtain the re-identification risk index of different medical records, and then generate dynamic masking rules for copying medical records based on all re-identification risk indices. The processing module is also used to perform desensitization processing on the medical record dataset through the dynamic masking rules, and determine the dynamic information availability rate of the medical record dataset after desensitization processing based on the completeness of the core content of different medical record records in the desensitization results. The execution module is used to authorize access to the batch copying of the medical record dataset based on the availability of the dynamic information and the preset permission verification rules.

[0013] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for desensitizing medical record data.

[0014] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for desensitizing medical record data.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The medical record copying and borrowing permission management system and method provided in this application obtains a medical record dataset for batch copying medical records during a research project. This dataset includes medical records of multiple patients and multiple fields within each record. The risk contribution of each field in patient identity re-identification is determined by the correlation strength between different fields in the dataset and their corresponding patient identities. Risk coupling is performed based on the risk contribution of all fields in patient identity re-identification, combined with field sensitivity, to obtain a re-identification risk index for different medical records. Dynamic masking rules for copying medical records are then generated based on all re-identification risk indices. The dataset is then de-identified using these dynamic masking rules. The dynamic information availability rate of the de-identified dataset is determined based on the completeness of the core content of different medical records in the de-identification result. Access authorization for batch copying of the medical record dataset is granted based on the dynamic information availability rate and preset permission verification rules.

[0016] Therefore, this solution, firstly, when determining the dynamic masking rules, distinguishes the risk contribution based on the association strength between different fields and patient identity, then combines field sensitivity for risk coupling. By quantifying the superposition effect of association strength and sensitivity, a re-identification risk index for each medical record is obtained, thereby generating dynamic masking rules that match the risk index. This process can accurately capture the differences in re-identification risk among different medical records, avoiding the problems of insufficient protection for high-risk medical records and excessive desensitization for low-risk medical records caused by traditional fixed desensitization rules. It provides a basis for risk quantification, enabling the desensitization rules to accurately match the actual risk level of the medical records. Then, in In determining the dynamic information availability rate: after desensitization using dynamic masking rules, the dynamic information availability rate is determined based on the completeness of the core content of the medical records in the desensitization results; this process can verify the impact of desensitization rules on the usability of research data in real time; this process provides a verification standard for the rationality of dynamic desensitization rules through quantitative feedback on availability rate, so that setting desensitization rules dynamically based on re-identification risk can not only protect privacy and security, but also avoid data loss of research value due to excessive desensitization, thus achieving a balance between risk control and usability; in summary, this scheme can dynamically set desensitization rules for medical record copying based on the re-identification risk of medical records. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a method for desensitizing medical record data according to some embodiments of this application; Figure 2 This is a schematic diagram illustrating the process of implementing desensitization according to some embodiments of this application; Figure 3 This is a schematic diagram illustrating the process of implementing access authorization according to some embodiments of this application; Figure 4 This is a schematic diagram of the structure of a medical record data desensitization unit according to some embodiments of this application; Figure 5 This is an internal structural diagram of a computer device for implementing a method for desensitizing medical record data, according to some embodiments of this application. Detailed Implementation

[0018] To better understand the technical solutions in this embodiment, the technical solutions in this embodiment will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0019] refer to Figure 1 The figure is a flowchart illustrating a method for desensitizing medical record data according to some embodiments of this application. The method 100 mainly includes the following steps: In step 101, a medical record dataset is obtained when medical records are copied in batches for a research project. The medical record dataset includes medical records of multiple patients and multiple fields in each medical record.

[0020] In practice, the bulk medical record data required for scientific research can be obtained through the medical institution's medical record management system. Preferably, structured medical record information (such as patient diagnosis and treatment records, treatment plans, and test results) can be extracted from the electronic medical record management system. Paper medical records that meet the scientific research inclusion criteria can be retrieved from the physical archive of medical records and digitized. The integrity of medical record data (such as supplementary medication records, imaging report conclusions, and test value trends) can be verified from the hospital information system. The authorized scope of data use (such as the permitted time period, disease type, and data fields) and privacy protection requirements can be extracted from the approval documents of the scientific research ethics review committee. The operation records of bulk copying (such as copying time, number of copies, and responsible person) can be extracted from the borrowing registration log of the medical records department. All of these can be integrated to form a medical record dataset that meets scientific research standards.

[0021] It should be noted that the medical record dataset for batch copying medical records in the research project mentioned in this application refers to a standardized collection of medical record information that supports the research topic and has been approved ethically. Physically, it is the sum of core medical information contained in a batch of medical records selected according to research inclusion and exclusion criteria (such as specific disease diagnosis, treatment time range, age limit, etc.). The selection of medical records is based on the "Regulations on Medical Record Management of Medical Institutions", the "Measures for Ethical Review of Medical Research", and the research project plan, and must simultaneously meet the following requirements: comply with the inclusion and exclusion criteria of the research design to ensure data relevance; be approved by the ethics review committee and obtain authorization to use the data; and de-identify patient information (such as deleting names, ID numbers, contact information, etc., and retaining only non-identification fields required for research) to protect patient privacy. The medical record refers to a standardized information carrier that records a patient's single or series of treatment processes and is the basic unit for research data collection.

[0022] In step 102, the risk contribution of each field in the patient identity re-identification is determined by the correlation strength between different fields in the medical record dataset and the corresponding patient identity.

[0023] In some embodiments, determining the risk contribution of each field in patient identity re-identification based on the association strength between different fields in the medical record dataset and the corresponding patient identity can be achieved through the following steps: Obtain unique identifiers for different patients in the medical record dataset; The strength of the association between different fields and corresponding patient identities in the medical record dataset is determined by all unique identifiers; Determine the re-identification risk characteristics of different fields in the medical record dataset; The risk contribution of each field in patient identity re-identification is determined based on all re-identification risk characteristics and all association strengths.

[0024] It should be noted that the unique patient identifier mentioned in this application refers to identification information used to distinguish different individual patients, which can achieve a unique mapping relationship of medical record data without directly revealing the patient's true identity. The unique identifier can be extracted from the patient's basic information fields (such as medical record number, hospital number, outpatient number, or internal code) recorded in the medical record management system; it can also be generated by encrypting or hashing fields containing identity features (such as name, ID number, contact information, etc.). The acquisition of the unique identifier can be performed during the medical record data preprocessing stage to perform normalization, desensitization, and duplicate value verification operations on the identity fields of each patient record, thereby ensuring... To ensure that the same patient corresponds to the same unique identifier in different medical records, in another embodiment, an identity association algorithm based on the patient's visit time sequence and the internal coding rules of the medical institution can be used to achieve unique identity mapping of medical records across departments or institutions. In some embodiments, the extraction frequency of the unique identifier can be set according to the update cycle of the medical record dataset. For example, the unique identifier can be generated or verified synchronously when the medical record is entered or modified daily to ensure identity consistency in subsequent risk calculations. In other embodiments, the automatic generation and updating of the unique identifier can also be achieved through a database trigger mechanism or dynamic synchronization of the identity index table. This application does not limit this.

[0025] In specific implementation, determining the association strength between different fields and corresponding patient identities in the medical record dataset through all unique identifiers can be achieved in the following way: First, for each field in the medical record dataset, obtain the specific value of that field in the medical record record corresponding to each unique patient identity identifier. Preferably, this includes first traversing all unique patient identity identifiers in the medical record dataset. Then, for each unique identifier, extract the value of the field in the medical record record corresponding to that unique identifier. The value includes the measurement result of a numerical field, the encoded representation of a text field, or the category label of a categorical field. Next, combine the values ​​of the field under all unique patient identities into a field-patient corresponding value set. Subsequently, for numerical fields, preferably, this includes first calculating the mean of the field in the field-patient corresponding value set; then calculating the deviation of each value from the mean. The covariance between the field and the patient's identity is obtained by summing the corresponding products. Next, the standard deviation of the field's value sequence is calculated, along with the standard deviation of the patient's unique identifier encoding sequence. Finally, the covariance is divided by the standard deviation product to obtain the Pearson correlation coefficient, which serves as the association strength between the numerical field and the patient's identity. Furthermore, for categorical fields, preferably, the process includes first uniformly encoding the field values ​​to ensure consistent numerical representation for each category in the calculation; then, statistically analyzing the probability distribution of each field value under different patient identities; next, calculating the mutual information value between the field value and the patient's identity based on this distribution; finally, using the mutual information value as the association strength between the categorical field and the patient's identity. In other embodiments, Spearman's rank correlation coefficient, conditional entropy, or information gain can also be used to calculate the association strength between the field and the patient's identity; this application does not limit this method.

[0026] It should be noted that the association strength mentioned in this application refers to the parameter value used to characterize the degree of statistical correlation between each field in the medical record dataset and the patient's identity. It is used to quantify the dependence of different fields on the patient's identity information in the dataset, thereby reflecting the potential privacy risks that the fields may pose to patient identification.

[0027] In specific implementation, determining the re-identification risk characteristics of different fields in the medical record dataset can be achieved in the following way: First, for each field in the medical record dataset, obtain its value sequence under the unique identifier of all patients; then, calculate the uniqueness of the field based on the value sequence. Preferably, this includes counting the number of different values ​​of the field and dividing by the total number of records to obtain the field uniqueness score. The higher the uniqueness, the more unique the field value, and the greater its risk contribution potential; next, calculate the stability of the field based on the value sequence. Preferably, this includes analyzing the degree of change of the field value over time, such as counting the consistency of the field value at different time points or in different medical record records. The higher the stability, the greater the risk contribution potential. The smaller the change in the value of the field, the stronger the sensitivity to patient identification. Subsequently, the uniqueness and stability are weighted and combined according to preset weights to generate the re-identification risk feature of the field. Preferably, a linear weighting method can be used, that is, the field re-identification risk feature is equal to the uniqueness multiplied by the uniqueness weight plus the stability multiplied by the stability weight, thereby obtaining a risk feature value normalized between 0 and 1, where 1 represents a high-risk feature and 0 represents a low-risk feature. In other embodiments, nonlinear weighting, entropy analysis, or information gain methods can also be used to comprehensively calculate uniqueness and stability. Preferably, the uniqueness weight and stability weight can be determined based on scientific research experience or the results of historical medical record analysis, which is not limited in this application.

[0028] It should be noted that the re-identification risk features mentioned in this application refer to parameter values ​​used to characterize the sensitivity of each field in the medical record dataset in potential identity recognition. They are used to quantify the risk that the uniqueness and stability of a field may cause to patient identity re-identification, thereby reflecting the importance of the field in privacy protection.

[0029] In specific implementation, the risk contribution of each field in patient identity re-identification can be determined based on all re-identification risk features and all association strengths in the following manner: First, for each field in the medical record dataset, obtain the re-identification risk feature value of the field and its association strength with patient identity; then, multiply the re-identification risk feature value of the field with its association strength to quantify the risk contribution of the field in patient identity re-identification; in other embodiments, nonlinear combination methods and weighted entropy values ​​can also be used to comprehensively calculate the re-identification risk features and the association strength to more comprehensively characterize the contribution of the field to patient identity re-identification, and this application does not limit this.

[0030] It should be noted that the risk contribution rate mentioned in this application refers to the numerical value of the importance parameter used to characterize the potential risk of each field in the medical record dataset in the process of patient identity re-identification. It is used to quantify the combined effect of the re-identification risk characteristics of a field and its association strength with patient identity, thereby reflecting the magnitude of the field's potential contribution to patient identity identification.

[0031] In step 103, risk coupling is performed based on the risk contribution of all fields in patient identity re-identification and field sensitivity to obtain the re-identification risk index of different medical records. Then, dynamic masking rules for copying medical records are generated based on all re-identification risk indices.

[0032] In some embodiments, risk coupling based on the risk contribution of all fields in patient identity re-identification and field sensitivity can be used to obtain the re-identification risk index for different medical records. This can be achieved through the following steps: Determine the field sensitivity of different fields in the medical record dataset; Select one medical record as the selected medical record; Retrieve all fields corresponding to the selected medical record; The risk contribution of all fields in patient identity re-identification and the field sensitivity of all fields are coupled into risk factors for each field; The re-identification risk index of selected medical records is determined based on all risk factors; Continue to determine the re-identification risk index for the remaining medical records.

[0033] In specific implementation, determining the sensitivity of different fields in the medical record dataset can be achieved in the following way: First, obtain the name and type information of all fields in the medical record dataset; then, classify the sensitivity of each field according to industry norms or relevant standards. Preferably, this includes classifying fields into three categories—high sensitivity, medium sensitivity, and low sensitivity—referring to the "Classification of Sensitive Information in Health and Medical Big Data." High sensitivity fields are those that can directly identify the patient or involve genetic privacy; medium sensitivity fields are those that require combination with other fields to identify the patient and are related to the core of diagnosis and treatment; and low sensitivity fields are those that cannot identify the patient and have no privacy attributes. Subsequently, determine the corresponding sensitivity coefficient for each sensitivity level. Preferably, the sensitivity coefficient for high sensitivity fields is 1.0, for medium sensitivity fields it is 0.6, and for low sensitivity fields it is 0.3. In other implementations, the sensitivity of fields can also be dynamically determined according to the hospital's internal privacy management regulations or the risk assessment opinions of the research ethics committee, or a sensitivity coefficient can be assigned to fields using expert scoring combined with statistical analysis results to ensure the accuracy and traceability of field sensitivity assessment. This application does not limit this approach.

[0034] It should be noted that the field sensitivity mentioned in this application refers to the parameter value used to characterize the degree to which each field in the medical record dataset involves patient identification or privacy information, in order to quantify the sensitivity of the field in data protection and privacy management, thereby guiding the formulation of de-identification strategies and access controls.

[0035] In specific implementation, coupling the risk contribution of all fields in patient re-identification and the field sensitivity of all fields into the risk factor of each field can be achieved in the following way: First, for each field in the medical record dataset, obtain the risk contribution value and sensitivity coefficient (i.e., field sensitivity) of that field; then, multiply the risk contribution value of that field with the corresponding sensitivity coefficient to quantify the comprehensive risk contribution of that field to patient re-identification in the medical record, and the resulting product is the risk factor of the corresponding field; in other embodiments, nonlinear combination methods, weighted averages, exponential amplification, or entropy analysis can also be used to couple the risk contribution and sensitivity coefficient to enhance the risk amplification effect and the ability to distinguish risk factors for highly sensitive fields, and this application does not limit this.

[0036] It should be noted that the risk factor mentioned in this application refers to a parameter value used to characterize the comprehensive risk contribution of each field in the medical record dataset in the process of patient identity re-identification. It is calculated by coupling the risk contribution of a field with the sensitivity of the field, and is used to quantify the potential privacy risks that a field may cause to patient identity recognition, thereby guiding the formulation of data desensitization and access control strategies.

[0037] In specific implementation, the re-identification risk index of a selected medical record can be determined based on all risk factors in the following manner: First, for the selected medical record, obtain the risk factors corresponding to all fields in the record; then, sum all risk factors to quantify the overall patient identity re-identification risk of the medical record, and the summation result is the re-identification risk index of the medical record; in other embodiments, a nonlinear mapping method can also be used to determine the re-identification risk index, which is not limited here.

[0038] It should be noted that the re-identification risk index mentioned in this application refers to a parameter value used to characterize the overall risk level of a single medical record in the process of patient identity re-identification. It is used to quantify the comprehensive privacy risk that the medical record may disclose patient identity information, thereby providing a basis for decision-making on dynamic desensitization and access control.

[0039] In some embodiments, the dynamic masking rules for generating photocopied medical records based on all re-identification risk indices can be implemented using the following steps: Set dynamic risk thresholds when photocopying medical records; The dynamic risk threshold and all re-identification risk indices are used to divide all medical records into sets of medical records with different risk levels. Dynamic masking rules for generating photocopied medical records based on sets of medical records with different risk levels.

[0040] It should be noted that the dynamic risk threshold mentioned in this application refers to the upper limit of the risk index used to distinguish the risk levels of different medical record records during the process of copying medical records. Its function is to guide the implementation of differentiated desensitization strategies for medical records with different risk levels, thereby balancing privacy protection and research usability. The dynamic risk threshold can be determined based on the ethical protection level of the research project, the sensitivity requirements of the medical records, and the data usage scenario. For example, the privacy protection level corresponding to the research project can be obtained from the ethical approval document or the hospital information system, and matched with a preset risk threshold range. The dynamic risk threshold can be represented by continuous or discrete numerical values, and can be dynamically set according to the update frequency of medical record data, the cycle of the research project, and the adjustment frequency of the desensitization strategy. For example, the threshold can be updated before each copying of medical records or during the monthly project review to ensure that the desensitization strategy of different batches of medical records is consistent with the latest risk assessment results. In other embodiments, statistical analysis methods or machine learning models can also be combined to automatically generate the dynamic risk threshold by fitting the distribution of the risk index of historical medical record re-identification. This application does not limit this.

[0041] In specific implementation, dividing all medical records into sets of medical records with different risk levels using the dynamic risk threshold and all re-identification risk indices can be achieved in the following way: First, for each record in the medical record dataset, obtain its corresponding re-identification risk index; then, compare the re-identification risk index with the dynamic risk threshold, and divide the medical records into a high-risk group and a normal-risk group based on the comparison result. Specifically, all medical records with a re-identification risk index greater than the dynamic risk threshold are classified as the high-risk group, and all medical records with a re-identification risk index lower than or equal to the dynamic risk threshold are classified as the normal-risk group. In other embodiments, a multi-threshold grading method can also be used to divide medical records into multiple risk levels, such as three or more levels (high, medium, low), to more finely reflect the re-identification risk level of medical records and form a corresponding risk level set. This application does not limit this approach.

[0042] It should be noted that the risk level medical record set mentioned in this application refers to the different groups formed by dividing medical records according to their risk level based on the comparison results of the re-identification risk index and dynamic risk threshold of a single medical record. This is used to distinguish between high-risk medical records and routine-risk medical records, or to divide them into multi-level risk groups, thereby reflecting the potential privacy risk level of each medical record in patient identity re-identification and providing a basis for dynamic desensitization and access control.

[0043] In specific implementation, the dynamic masking rules for generating medical record copies based on medical record sets of different risk levels can be implemented in the following way: First, for each risk level of medical record set, obtain the field information of all medical record records in the set and the risk contribution of each field in patient identity re-identification; then, determine the desensitization strategy based on the field risk contribution and its risk level. Preferably, fields with higher risk contribution in the high-risk group are subjected to strong desensitization operations, such as hashing, masking some characters, or retaining only the year information; fields with higher risk contribution in the normal risk group can be subjected to weak desensitization operations, such as retaining some information or performing interval processing; low-risk fields can choose to directly retain the original value, thereby generating dynamic masking rules for generating medical record copies. In other implementations, dynamic masking rules can also be generated through a rule engine or automated script, combined with research needs, data availability, and privacy protection requirements.

[0044] It should be noted that the dynamic masking rule described in this application refers to the operational strategy used to quantify and control the patient identification identifiability of each field in medical record data under different risk levels. It is a mapping relationship generated by coupling the risk contribution of a field with the risk level of the medical record. This application uses this mapping relationship to determine the type and intensity of desensitization operation to be performed on each field when copying medical records, thereby achieving dynamic control of the risk of patient identification.

[0045] In step 104, the medical record dataset is de-identified using the dynamic masking rules, and the dynamic information availability rate of the medical record dataset after de-identification is determined based on the completeness of the core content of different medical record records in the de-identification results.

[0046] In some embodiments, reference Figure 2 As shown in the figure, this is a schematic diagram illustrating the process of de-identification processing in some embodiments of this application. The de-identification processing of the medical record dataset using the dynamic masking rules can be achieved through the following steps: First, in step 1041, the risk levels corresponding to different medical records in the medical record dataset are obtained; Then, in step 1042, an adaptive masking scheme for each medical record is generated based on the dynamic masking rules and the risk levels corresponding to different medical records. Finally, in step 1043, the medical record dataset is desensitized according to the adaptive masking scheme of each medical record.

[0047] In specific implementation, the adaptive masking scheme for each medical record, generated through the dynamic masking rules and the risk levels corresponding to different medical record records, can be implemented in the following way: First, for each record in the medical record dataset, obtain its corresponding risk level information; then, based on the risk level, traverse all fields in the medical record field by field, and for each field, determine its corresponding de-identification operation in the dynamic masking rules, specifically including: obtaining the field type (high-risk field, medium-risk field, or low-risk field), obtaining the de-identification strategy for that type under the risk level (such as hashing, partial information preservation, etc.). (The data can be retained, ranged, or the original value can be retained), and the field name, field type, risk level, and corresponding desensitization operation can be combined into a field desensitization mapping record; then, the desensitization mapping records of all fields in the record are summarized to form an adaptive masking scheme for the medical record; in other embodiments, all medical records can also be processed in batches through a rule engine or automated script to achieve dynamic matching of each record field with the desensitization operation, so as to ensure that high-risk records perform strong desensitization and regular risk records perform weak desensitization, and the desensitization scheme can be dynamically updated according to research needs or project cycle, which is not limited in this application.

[0048] It should be noted that the adaptive masking scheme described in this application refers to a set of strategies for dynamically allocating field desensitization operations for each medical record based on its risk level. It is a mapping result generated by matching the risk level information of the medical record with the desensitization operation mapping relationship of the fields in the dynamic masking rules. This application uses this mapping result to determine the specific desensitization operation type and intensity that should be performed on each field in each record during copying or use, thereby realizing personalized and dynamic desensitization control for medical records with different risk levels.

[0049] In specific implementation, the anonymization of the medical record dataset based on the adaptive masking scheme of each medical record can be achieved in the following way: First, for each medical record in the medical record dataset, obtain its corresponding adaptive masking scheme; the adaptive masking scheme includes the field type, risk level, and corresponding anonymization operation of each field in the medical record; then, traverse the medical record field by field, and process the original data according to the anonymization operation corresponding to each field in the adaptive masking scheme. The anonymization operation includes, but is not limited to: hashing (e.g., performing SHA256 hashing on the ID number field), partial information retention (e.g., retaining only the year or month for the birth date field), interval processing (e.g., mapping numeric fields to a specified interval range), and retaining the original value (e.g., keeping the department name field unchanged); subsequently, summarize the anonymized field data to form the anonymized record of the medical record; the operation can be performed sequentially on all records in the medical record dataset to generate a complete anonymized medical record dataset as the anonymization result.

[0050] In some embodiments, determining the dynamic information availability of the medical record dataset after desensitization based on the completeness of the core content of different medical record records in the desensitization results can be achieved through the following steps: Obtain a list of core information corresponding to the research project; The completeness of the core content of different medical record entries in the desensitization results is determined by the core information list; The availability of the anonymized medical record dataset is verified based on the completeness of the core content of different medical record records, thus obtaining the dynamic information availability rate of the anonymized medical record dataset.

[0051] It should be noted that the core information list of the research project mentioned in this application refers to the key fields and their content requirements that need to be retained in the medical record data to ensure the smooth conduct of research analysis. This list can be extracted from the research project application, ethics approval documents, or hospital information system. The core information list includes, but is not limited to, core fields such as diagnosis code, generic name of medication, key experimental indicators, and desensitized date of birth. The core information list can be determined according to the research objectives, data analysis needs, and privacy protection requirements of the research project. For example, for the project "Efficacy Analysis of Medication for Type 2 Diabetes", the core fields include diagnosis code (subtype needs to be distinguished), generic name of medication (complete), glycated hemoglobin value (range needs to be defined), and desensitized date of birth (age group analysis needs to be supported).

[0052] In specific implementation, determining the completeness of the core content of different medical records in the desensitized results through the core information list can be achieved in the following way: First, for each medical record in the desensitized medical record dataset, obtain all the core fields listed in the research project's core information list in that record in sequence; then, for each core field, compare it according to the content requirements specified in the list, specifically including: comparing whether the number of digits retained or the code prefix of the desensitized diagnostic code meets the research analysis's requirements for distinguishing disease subtypes; for example, after desensitization, only the first three digits of the code are retained for the high-risk group, and its completeness is determined by checking whether the first three digits can still distinguish the main disease; checking whether the desensitized field is missing or replaced, confirming whether the field content is a complete drug name, and determining incompleteness if the field is empty or only contains partial characters; comparing the desensitized numerical or regional values. Whether the granularity meets the minimum requirements for scientific research analysis, such as whether the interval range is fine enough for efficacy grading; whether the date information after anonymization can be mapped to the age range required for scientific research analysis; preferably, for the comparison of each core field, a rule engine, condition judgment or automated script can be used to execute, for example: checking whether the field is empty, whether the value is within the allowed range, whether the string length meets the minimum requirements; generating Boolean values ​​or status markers for the comparison results; subsequently, by statistically analyzing the proportion of core fields that meet the scientific research requirements (i.e., judged to be complete) in the medical record to the total number of core fields, the completeness of the core content of the corresponding medical record is taken. In other embodiments, the comparison results can also be weighted according to the importance level of the core fields to form a more refined state of core content completeness. This application does not limit this.

[0053] It should be noted that the core content completeness mentioned in this application refers to the parameter value used to quantify the completeness of the fields in the medical record after anonymization that are included in the core information list of the scientific research project in terms of meeting the requirements of scientific research analysis content.

[0054] In specific implementation, the usability of the anonymized medical record dataset is verified based on the completeness of the core content of different medical record records. The dynamic information usability rate of the anonymized medical record dataset can be achieved in the following way: First, the completeness value of the core content corresponding to each medical record in the anonymized medical record dataset is obtained sequentially; then, the completeness value of the core content of each medical record is compared with the minimum completeness requirement preset by the research project to determine whether the medical record can be used for research analysis; wherein, the minimum completeness requirement can be set according to the analytical needs of the research project and the importance of the core fields, specifically including: classifying the importance of all core fields, assigning weight coefficients to each, and determining the importance of the core fields in research analysis. The necessity of determining the minimum completeness value is as follows: For example, when a research project requires that all three types of fields—diagnosis, medication, and key experimental indicators—are present, the minimum completeness requirement can be set to 80% of the total number of core fields. During the comparison process, by judging whether the completeness value of the core content of the medical record is greater than or equal to the minimum completeness requirement, it is determined whether the medical record meets the minimum completeness standard required for scientific research analysis. If it meets the standard, it is marked as a usable record; if it does not meet the standard, it is marked as an unusable record. Then, the number of all records marked as usable is counted, and the total number of records in the anonymized medical record dataset is obtained. Next, the ratio of the number of usable records to the total number of records is calculated, and the determined result is used as the dynamic information availability rate of the medical record dataset.

[0055] It should be noted that the dynamic information availability rate mentioned in this application refers to a parameter value used to quantify the proportion of records in the anonymized medical record dataset that can be used for scientific research analysis, in order to reflect the overall availability level of the anonymized medical record dataset for scientific research analysis while ensuring privacy.

[0056] In step 105, access authorization is granted for batch copying of the medical record dataset based on the dynamic information availability rate and preset permission verification rules.

[0057] In some embodiments, granting access authorization for batch copying of the medical record dataset based on the dynamic information availability rate and preset permission verification rules can be achieved through the following steps: The desensitization process of the medical record dataset is dynamically adjusted based on the availability of the dynamic information. Based on preset permission verification rules and adjustment results, a permission certificate for batch copying of the medical record dataset is generated; Access is granted for batch copying of the medical record dataset based on the aforementioned permission credentials.

[0058] In specific implementation, the dynamic granularity adjustment of the desensitization process of the medical record dataset based on the dynamic information availability rate can be achieved in the following way: First, obtain the dynamic information availability rate of the desensitized medical record dataset and compare the availability rate with a preset availability rate level range to determine the current availability status of the medical record dataset; then, based on the determined availability status, select the corresponding desensitization granularity template from the desensitization strategy rule base; then, adjust the desensitization granularity of the core fields in the medical record dataset according to the desensitization granularity template. When the availability status is low, automatically select a strategy to reduce the desensitization intensity, such as extending the diagnostic code retention bits. The experiment index grading range is adjusted by increasing the number or relaxing the range. When the availability status is high, the desensitization intensity is increased, such as by increasing the masking ratio or blurring the accuracy of birth dates. Then, based on the desensitization granularity adjustment result, the dynamic information availability rate of the medical record dataset is updated and compared with the target availability rate threshold. When the difference is less than a preset convergence range, the result of this round of desensitization granularity adjustment is taken as the final adjustment result. It should be noted that the availability rate range is a set of preset intervals based on research needs and privacy protection standards. For example, when the availability rate is below 80%, it is considered low availability, indicating excessive data desensitization; when the availability rate is between 80% and 90%, it is considered high availability. The availability balance state indicates that the data has reached a balance between privacy and scientific research availability; when the availability rate is higher than 90%, it is a high availability state, indicating that the de-identification intensity is low; in addition, the range value can be determined by the data availability statistics of historical scientific research projects or by simulation test based on privacy risk model; the de-identification strategy rule base consists of multiple de-identification strategy units, each strategy unit includes field identifier, field sensitivity level, availability target value and corresponding de-identification parameter group (specifically including the number of bits to retain, mask type, degree of obfuscation, etc.).

[0059] In specific implementation, the generation of authorization credentials for batch copying of the medical record dataset based on preset authorization verification rules and adjustment results can be achieved in the following way: First, obtain the desensitized medical record dataset after dynamic granular adjustment (i.e., the adjustment result) and the corresponding dynamic information availability rate, and obtain the applicant's identity information and research project application content; then, compare the dynamic information availability rate with a preset minimum availability rate threshold. When the availability rate reaches the minimum availability rate threshold, the authorization verification process is initiated; otherwise, authorization credentials are not generated; then, the applicant is verified layer by layer according to the preset authorization verification rules, including identity qualification verification, application requirement matching verification, and ethical compliance verification. Among these, identity qualification verification includes checking whether the applicant is on the research project filing list and whether the institution has the qualifications to use medical data. The validity of the application is checked; the application requirement matching verification includes checking whether the scope and fields of the requested copy are consistent with the core information list of the research project; the ethical compliance verification includes checking whether the application is within the scope of ethical approval and whether the approval validity period has not expired; then, after all the above verifications are passed, the hospital's access control system is called to generate a structured access certificate, which includes a unique certificate number, research project information, authorized copying record scope and fields, validity period and security and anti-counterfeiting information, etc.; as a preferred embodiment, the generation of the access certificate can adopt a key-value mapping method, mapping the certificate number to the authorized access medical record and field list, so that the system can quickly verify and control the batch copying operation. In other embodiments, digital signatures or blockchain methods can be used to store the certificate information to enhance anti-tampering and traceability, which is not limited here.

[0060] It should be noted that the authorization credentials mentioned in this application refer to structured credential information used to authorize and control bulk copying access to medical record datasets. This information is used to record the scope, fields, validity period, and security and anti-counterfeiting information of the authorized medical records to be copied, thereby ensuring that the medical record copying operation is traceable, verifiable, and dynamically manageable under the premise of meeting privacy protection and scientific research compliance requirements.

[0061] For specific implementation, refer to Figure 3As shown in the figure, this is a schematic diagram of the access authorization process in some embodiments of this application. Access authorization for batch copying of the medical record dataset based on the permission credentials can be implemented in the following way: First, the permission credentials are associated with the de-identified medical record dataset in the system. This includes assigning a unique credential number to each permission credential and establishing a mapping relationship between the credential number and the accessible medical record record number and field list in the database or secure storage. Then, the applicant's identity is confirmed to be consistent with the credential holder through identity verification methods, preferably including facial recognition, fingerprint recognition, terminal authorization code, or dynamic password verification. Next, on the copying terminal... After scanning the credential number, the system verifies whether the applicant has the right to access the specified medical record and fields based on the mapping relationship, and checks whether the credential is valid and complies with terminal or network access restrictions. When all verifications pass, the applicant is allowed to perform batch copying operations within the authorized scope, outputting only the de-identified core field information, thereby enabling access authorization for batch copying of the medical record dataset. As a preferred embodiment, the system can store the mapping relationship in the form of key-value mapping or hash table to quickly verify and control access permissions. In other embodiments, the permission credential and mapping relationship can also be stored in the form of digital signature or blockchain to enhance anti-tampering and traceability. This application does not limit this.

[0062] In another aspect, in some embodiments, this application provides a medical record copying and borrowing permission management system, including a medical record data desensitization unit, as referenced. Figure 4 The figure is a schematic diagram of the structure of a medical record data desensitization unit according to some embodiments of this application. The medical record data desensitization unit 200 includes: an acquisition module 201, a processing module 202, and an execution module 203, which are described below: The acquisition module 201 in this application is mainly used to acquire the medical record dataset when batch copying medical records for scientific research projects. The medical record dataset includes medical records of multiple patients and multiple fields in each medical record. Processing module 202, in this application, is mainly used to determine the risk contribution of each field in the patient identity re-identification by the correlation strength between different fields in the medical record dataset and the corresponding patient identity; In addition, the processing module 202 in this application is also used to perform risk coupling based on the risk contribution of all fields in patient identity re-identification and the field sensitivity to obtain the re-identification risk index of different medical records, and then generate dynamic masking rules for copying medical records based on all re-identification risk indices. In addition, the processing module 202 in this application is also used to perform desensitization processing on the medical record dataset through the dynamic masking rules, and determine the dynamic information availability rate of the medical record dataset after desensitization processing based on the completeness of the core content of different medical record records in the desensitization results. The execution module 203 in this application is mainly used to authorize access to the batch copying of the medical record dataset based on the availability of the dynamic information and the preset permission verification rules.

[0063] In addition, this application also provides a computer device, the computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described method for desensitizing medical record data.

[0064] In some embodiments, reference Figure 5 The figure is an internal structural diagram of a computer device implementing a medical record data desensitization method according to some embodiments of this application. The medical record data desensitization method in the above embodiments can be implemented through... Figure 5 The computer device shown is used to implement this, and the computer device 300 includes at least one processor 301, a communication bus 302, a memory 303, and at least one communication interface 304.

[0065] The processor 301 may be a general-purpose central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more devices used to control the execution of the medical record data desensitization method in this application.

[0066] The communication bus 302 is used to transmit information between the aforementioned components.

[0067] Memory 303 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CDROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory 303 may exist independently and be connected to processor 301 via communication bus 302. Memory 303 may also be integrated with processor 301.

[0068] The memory 303 stores program code for executing the scheme of this application, and its execution is controlled by the processor 301. The processor 301 executes the program code stored in the memory 303. The program code may include one or more software modules. In the above embodiments, the medical record data desensitization method can be implemented by the processor 301 and one or more software modules in the program code in the memory 303.

[0069] Communication interface 304 uses any transceiver-like device to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0070] In a specific implementation, as one example, a computer device may include multiple processors, each of which may be a single-core processor or a multi-core processor. Here, a processor may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0071] The aforementioned computer device can be a general-purpose computer device or a special-purpose computer device. In specific implementations, the computer device may be a desktop computer, a portable computer, a network server, a handheld digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. This application does not limit the type of computer device.

[0072] In addition, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for desensitizing medical record data.

[0073] In summary, the medical record copying and borrowing permission management system and method disclosed in this application obtains a medical record dataset for batch copying of medical records during research projects. This dataset includes medical records of multiple patients and multiple fields within each record. The risk contribution of each field in patient identity re-identification is determined by the correlation strength between different fields in the dataset and their corresponding patient identities. Risk coupling is performed based on the risk contribution of all fields in patient identity re-identification, combined with field sensitivity, to obtain a re-identification risk index for different medical records. Dynamic masking rules for copying medical records are then generated based on all re-identification risk indices. The dataset is de-identified using these dynamic masking rules, and the dynamic information availability rate of the de-identified dataset is determined based on the completeness of the core content of different medical records in the de-identification result. Access authorization for batch copying of the medical record dataset is granted based on the dynamic information availability rate and preset permission verification rules. De-identification rules for copying medical records can be dynamically set based on the re-identification risk of the medical records.

[0074] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0075] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for de-identifying medical record data, used in a medical record copying and borrowing access management system to authorize copying access to de-identified medical record data, characterized in that, The method includes the following steps: Obtain the medical record dataset when batch copying medical records for scientific research projects. The medical record dataset includes medical records of multiple patients and multiple fields in each medical record. The risk contribution of each field in patient identity re-identification is determined by the correlation strength between different fields in the medical record dataset and the corresponding patient identities. Risk coupling is performed based on the risk contribution of all fields in patient identity re-identification and field sensitivity to obtain the re-identification risk index of different medical records. Then, dynamic masking rules for copying medical records are generated based on all re-identification risk indices. The medical record dataset is anonymized using the dynamic masking rules, and the dynamic information availability rate of the medical record dataset after anonymization is determined based on the completeness of the core content of different medical record records in the anonymization results. Access authorization is granted for batch copying of the medical record dataset based on the availability of the dynamic information and the preset permission verification rules.

2. The method as described in claim 1, characterized in that, Determining the risk contribution of each field in patient identity re-identification by the association strength between different fields in the medical record dataset and their corresponding patient identities specifically includes: Obtain unique identifiers for different patients in the medical record dataset; The strength of the association between different fields and corresponding patient identities in the medical record dataset is determined by all unique identifiers; Determine the re-identification risk characteristics of different fields in the medical record dataset; The risk contribution of each field in patient identity re-identification is determined based on all re-identification risk characteristics and all association strengths.

3. The method as described in claim 1, characterized in that, Based on the risk contribution of all fields in patient identity re-identification and combined with field sensitivity, a risk coupling is performed to obtain the re-identification risk index for different medical records, which specifically includes: Determine the field sensitivity of different fields in the medical record dataset; Select one medical record as the selected medical record; Retrieve all fields corresponding to the selected medical record; The risk contribution of all fields in patient identity re-identification and the field sensitivity of all fields are coupled into risk factors for each field; The re-identification risk index of selected medical records is determined based on all risk factors; Continue to determine the re-identification risk index for the remaining medical records.

4. The method as described in claim 1, characterized in that, The dynamic masking rules for generating photocopied medical records based on all re-identification risk indices specifically include: Set dynamic risk thresholds when photocopying medical records; The dynamic risk threshold and all re-identification risk indices are used to divide all medical records into sets of medical records with different risk levels. Dynamic masking rules for generating photocopied medical records based on sets of medical records with different risk levels.

5. The method as described in claim 1, characterized in that, The desensitization process for the medical record dataset using the dynamic masking rules specifically includes: Obtain the risk level corresponding to different medical records in the medical record dataset; An adaptive masking scheme is generated for each medical record based on the dynamic masking rules and the risk levels corresponding to different medical records. The medical record dataset is anonymized based on an adaptive masking scheme for each medical record.

6. The method as described in claim 1, characterized in that, The dynamic information availability rate of the medical record dataset after desensitization is determined based on the completeness of the core content of different medical record records in the desensitization results. Specifically, this includes: Obtain a list of core information corresponding to the research project; The completeness of the core content of different medical record entries in the desensitization results is determined by the core information list; The availability of the anonymized medical record dataset is verified based on the completeness of the core content of different medical record records, thus obtaining the dynamic information availability rate of the anonymized medical record dataset.

7. The method as described in claim 1, characterized in that, Access authorization for batch copying of the medical record dataset based on the dynamic information availability rate and preset permission verification rules specifically includes: The desensitization process of the medical record dataset is dynamically adjusted based on the availability of the dynamic information. Based on preset permission verification rules and adjustment results, a permission certificate for batch copying of the medical record dataset is generated; Access is authorized for batch copying of the medical record dataset based on the aforementioned permission credentials.

8. A medical record photocopying and borrowing access control system, characterized in that, It includes a medical record data desensitization unit, which includes: The acquisition module is used to acquire the medical record dataset when batch copying medical records for scientific research projects. The medical record dataset includes medical records of multiple patients and multiple fields in each medical record. The processing module is used to determine the risk contribution of each field in the patient identity re-identification process by the correlation strength between different fields in the medical record dataset and the corresponding patient identity. The processing module is also used to perform risk coupling based on the risk contribution of all fields in patient identity re-identification and the field sensitivity to obtain the re-identification risk index of different medical records, and then generate dynamic masking rules for copying medical records based on all re-identification risk indices. The processing module is also used to perform desensitization processing on the medical record dataset through the dynamic masking rules, and determine the dynamic information availability rate of the medical record dataset after desensitization processing based on the completeness of the core content of different medical record records in the desensitization results. The execution module is used to authorize access to the batch copying of the medical record dataset based on the availability of the dynamic information and preset permission verification rules.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the medical record data desensitization method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the medical record data desensitization method as described in any one of claims 1 to 7.