Hospital database management method and system based on adaptive data desensitization

By adopting adaptive data desensitization methods in hospital databases, dynamic grading and desensitizing multi-source data, combined with reinforcement learning and abnormal detection technology, the problem of imbalance between data security and availability of traditional methods is solved, and efficient data utilization and secure sharing are achieved.

CN120030601AActive Publication Date: 2025-05-23THE 900TH HOSPITAL OF THE CHINESE PEOPLES LIBERATION ARMY JOINT LOGISTICS SUPPORT FORCE

Patent Information

Application Number
CN202510511995.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Traditional static data desensitization methods are difficult to adapt to complex hospital scenarios, and are prone to imbalance between data security and availability, which cannot effectively meet the needs of hospitals in data protection and privacy compliance.

Method used

Adopting hospital database management method based on adaptive data desensitization is adopted, and multi-source data is hierarchical through data classification algorithms, and a desensitization model based on role permissions, access content and data characteristics is established, and optimal desensitization strategies are dynamically generated, and abnormal detection is performed through reinforcement learning and GANomaly model.

Benefits of technology

It realizes efficient use and sharing of data while ensuring data security, improves the intelligence level of hospital data management, and can dynamically adjust the desensitization strategy according to user roles and access scenarios, improving the accuracy and security of data desensitization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030601A_ABST
    Figure CN120030601A_ABST
Patent Text Reader

Abstract

The invention relates to a hospital database management method and system based on adaptive data desensitization. The method comprises the following steps: S1, uniformly integrating multi-source data of a hospital into a database; s2, performing sensitivity grading on the data of the database through a data classification algorithm, and automatically generating corresponding labels; s3, establishing a desensitization model based on role permission, access content and data features; s4, the user inquires data and inputs related parameters through the front-end system, whether fields or data in the request contain high-sensitivity fields or not is judged, and if the fields or the data contain the high-sensitivity fields, a proper desensitization strategy is dynamically generated based on a desensitization model according to the user role, the data sensitivity and the access context; and S5, monitoring access behaviors in combination with data using log records, carrying out anomaly detection based on a GANopen model, and if an abnormal behavior is detected, triggering an alarm or stopping access. According to the invention, efficient utilization and sharing of data are realized, and the intelligent level of hospital data management is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data management, and in particular to a hospital database management method and system based on adaptive data desensitization. Background Art

[0002] With the rapid development of medical information technology, hospitals have accumulated a large amount of sensitive data containing patient privacy in their daily operations. These data mainly come from electronic medical record systems (EMR), laboratory information systems (LIS), picture archiving systems (PACS), and hospital information management systems (HIS). These data provide support for hospital diagnosis, treatment, scientific research, and operations, but their highly sensitive attributes (such as patient names, ID numbers, medical records, etc.) expose them to high risks of data leakage, abuse, and illegal access. At the same time, in the context of big data analysis, artificial intelligence applications, and testing, the need for data openness further increases the possibility of data leakage.

[0003] With the increasing attention paid to data sensitivity, higher requirements are placed on data management: not only is it necessary to ensure that sensitive data is not leaked, but it is also necessary to meet diverse usage requirements on the basis of compliance. However, traditional static data desensitization methods are difficult to adapt to the complex scenarios of hospitals and are prone to imbalance between security and availability. Summary of the invention

[0004] In order to solve the above problems, the purpose of the present invention is to provide a hospital database management method and system based on adaptive data desensitization, which can not only meet the hospital's needs in data protection and privacy compliance, but also achieve efficient use and sharing of data while ensuring data security, thereby improving the intelligence level of hospital data management.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A hospital database management method based on adaptive data desensitization comprises the following steps: S1: Integrate the hospital's multi-source data into a unified database; S2: Classify the sensitivity of the multi-source data in the database through a data classification algorithm and automatically generate corresponding labels; S3: Adjust the desensitization strategy in real time through context analysis and permission matching, establish a desensitization model based on role permissions, access content and data features, take user requests as input, and output the optimal combination of data desensitization rules; S4: The user queries data through the front-end system, enters relevant parameters, and determines whether the fields or data in the request contain highly sensitive fields. If they do, a suitable desensitization strategy is dynamically generated based on the desensitization model according to the user role, data sensitivity, and access context; S5: Combine data usage logs to monitor access behavior and perform anomaly detection based on the GANomaly model. If abnormal behavior is detected, a warning will be triggered or access will be blocked.

[0006] Furthermore, the hospital's multi-source data is integrated into the database, as follows: The hospital's multi-source data includes the electronic medical record system EMR, the laboratory information system LIS, the image archiving system PACS and the operation management system HIS; Among them, EMR stores patient electronic medical records, including text data and structured data; LIS stores laboratory test results, which are structured table data; PACS stores medical images, and the data is mainly unstructured files, and also contains relevant metadata; HIS stores hospital operation data, including financial systems and medical records, and the data is in the form of structured tables; The HL7 protocol is used to standardize the interfaces of different data systems, and unstructured data is associated with structured data through metadata or indexing mechanisms; finally, the integrated data is stored in an integrated database.

[0007] Furthermore, the sensitivity of the data in the database is graded through a data classification algorithm, and corresponding labels are automatically generated, as follows: Extract all tables, fields and record data from the database, including field names, field values, and table context information; Using regular matching, keyword recognition, and TF-IDF / NLP models, we calculate the comprehensive sensitivity score S of each field and table. The sensitivity levels are divided according to the thresholds, and labels are automatically generated and written into the meta-information table of the database.

[0008] Furthermore, the comprehensive sensitivity score S is the weighted result of multiple characteristic indicators, as follows: Use predefined regular rules to match field names and field values ​​to determine whether they contain common personal privacy information. Use the NLP model to perform semantic analysis on field names and context content to determine whether the field semantically belongs to personal identity information (PII). Based on the results of rule judgment and semantic analysis, assign a score S. PII ; ; in, is the embedding vector of the field name; is a predefined PII classification template vector; It is the result of field name rule detection; is the result of semantic analysis; Analyze the content characteristics of fields or records, identify whether the field value contains privacy information or sensitive keywords, and calculate the content sensitivity score S content : ; Among them, W k is the keyword weight; Count(k,C i ) is the field value C i Total_Words(C i ) is the field value C i The total number of words; K is the total number of keywords; Combine the field name and the context semantics of the table to determine the sensitivity of the field content and obtain the context sensitivity score S context ; ; in, is the business sensitive weight; is the similarity between the field name and the sensitive semantic template; Determine the uniqueness and sensitivity of the field value based on its distribution and obtain the uniqueness score S uniqueness ; ; Among them, Unique_Count(F i ) is F i Total_Records(F i ) is the total number of records; Based on S PII , S content , S context and S uniqueness Calculate the comprehensive sensitivity score S total : ; in, is the weight coefficient.

[0009] Furthermore, the desensitization model models the process of desensitizing user requests and data sensitive information as a multi-dimensional optimization problem. The goal is to formulate the global optimal desensitization strategy combination through dynamic analysis of user context, data features and access context. At the same time, the model outputs the optimal desensitization rules and desensitized data, as follows: Define each F i , calculate F i The corresponding sensitivity score S i, user role R and the corresponding permission level Access_Threshold(R); access purpose M is the context; context adjustment factor C for sensitivity score m , used to dynamically relax or tighten sensitivity constraints; The model goal is to select the embedding vector F of the field name i Desensitization strategy , the final optimization goal is: ; ; in, Score the security of the policy in the current context; The ability to preserve data availability for policy; The desensitization cost of the strategy; is the weight coefficient; To optimize the objective function.

[0010] Furthermore, the desensitization model uses reinforcement learning DQL to continuously optimize desensitization rules and improve strategies based on user behavior data, as follows: The current state of the system S={S user ,S data ,S context} represents the context of the desensitization question, where S user is the user context; S data is the data context; S context To access the environment context; The set of desensitization strategies selected for each data field is A. Each action corresponds to a desensitization method. Reinforcement learning DQL will select an optimal action for each field. By balancing data security and data availability, rewards That is, the optimization goal of the desensitization model: ; Among them, R safety Indicates whether data security requirements are met. If the field sensitivity score S i If it is greater than the user permission threshold and is not desensitized, ; If desensitization is correct, then R safety =1; R utility Whether the desensitized data still has statistical or business value is defined as the data availability score after desensitization; R complexity represents the cost of executing different desensitization strategies; is the weight coefficient; The strategy π represents the probability distribution of the action selected by the system each time: π(A|S;θ); Where θ is the parameter fitted by the deep neural network; Reinforcement learning DQL guides policy optimization through the value function Q(S,A): ; Among them, S' is the new state reached after the current action; A' is the action set corresponding to the new state; is the reward discount factor; Use a neural network to fit the Q function, input the state S, and output the Q(S,A) value of each action.

[0011] Furthermore, the training process of the neural network is as follows: Record historical status, actions, rewards, and state transfers ; Randomly extract samples from the experience replay pool to train the network and reduce sample correlation; The training goal is to minimize the error between the value function and the DQN's current estimate: ; ; in, is the target network parameter; represents the expected value of the sample (S, A, R, S′) in the experience playback sampling data; Y represents the target Q value; Indicates that the target network parameters are Q value; Indicates that the target network parameters are Q value; is the objective function; By training DQN, the optimal action is found in each state S : .

[0012] Further, S4 is specifically: Enter user role R, access data field F i , adjustment factor C m , and by adjusting the factor C m Adjust the sensitivity score: ; Among them, S i ′Adjusted sensitivity score; Check Action Whether the permission is satisfied, if not, it returns no access permission; if satisfied, the desensitization strategy utility calculation is performed: Traverse T i ={Ti1 ,T i2 ,...,T ij, ...,T iJ}, calculate the utility of each strategy: ; For the utility score U(T ij ,F i ) takes the maximum value and selects the optimal strategy: ; Apply the optimal strategy to output the desensitized data.

[0013] Furthermore, anomaly detection is performed based on the GANomaly model. Based on the structure of the generative adversarial network GAN, reconstruction error, latent space consistency, and the discriminant ability of the discriminator are introduced to calculate the anomaly score, as follows: The encoder E maps the original input behavior feature X to a latent space vector z: ; in, The trainable parameters of the encoder; z is a latent space vector, which is a low-dimensional compressed representation of the behavior feature X; The generator G reconstructs the approximate features based on the latent space vector z generated by the encoder : ; in, is the trainable parameter of the generator; The discriminator D is responsible for distinguishing the input behavior features X from the features generated by the generator. , output a probability value D(X)∈[0,1] to determine whether the input is real data or generated reconstructed data:

[0014] in, is the activation function, is the weight matrix of the discriminator, is bias; In GANomaly, the loss function Loss due to reconstruction error , latent space consistency loss and the discriminative loss composition: ; ; ; ; Among them, λ 1 ,λ 2 is the weight coefficient, Features generated for the generator The output after inputting the encoder, Features generated for the generator The output after inputting the discriminator, is the expected function.

[0015] A hospital database management system based on adaptive data desensitization includes a processor, a memory, and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the hospital database management method based on adaptive data desensitization as described above.

[0016] The present invention has the following beneficial effects: 1. The present invention integrates PII identification, content sensitivity analysis, contextual semantic analysis and uniqueness detection through a sensitivity grading model, automatically divides hospital database data into highly sensitive, medium sensitive and low sensitive levels, and generates traceable sensitivity labels, providing a technical basis for data security, dynamic desensitization and compliance management; 2. The present invention can dynamically generate the optimal desensitization strategy according to user roles, access content and scenarios (such as diagnosis, treatment, scientific research, and testing), avoiding the "one-size-fits-all" problem of static desensitization rules; and continuously optimize desensitization rules through reinforcement learning (such as Deep Q-Learning), and can continuously learn and improve strategies based on user behavior data, making the desensitization effect more accurate; 3. The present invention is based on the GANomaly model, through the collaborative work of the encoder, generator and discriminator, using normal behavior modeling as a baseline to capture abnormal behaviors that deviate from the normal pattern. By reconstructing errors and latent space consistency, it can accurately detect abnormal data queries, unauthorized behaviors or batch sensitive data access, thereby protecting data security. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0018] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments: refer to Figure 1 In this embodiment, a hospital database management method based on adaptive data desensitization is provided, comprising the following steps: S1: Integrate the hospital's multi-source data into a unified database; S2: Classify the sensitivity of the multi-source data in the database through a data classification algorithm and automatically generate corresponding labels; S3: Adjust the desensitization strategy in real time through context analysis and permission matching, establish a desensitization model based on role permissions, access content and data features, take user requests as input, and output the optimal combination of data desensitization rules; S4: The user queries data through the front-end system, enters relevant parameters, and determines whether the fields or data in the request contain highly sensitive fields. If they do, a suitable desensitization strategy is dynamically generated based on the desensitization model according to the user role, data sensitivity, and access context; S5: Combine data usage logs to monitor access behavior and perform anomaly detection based on the GANomaly model. If abnormal behavior (such as unauthorized queries) is detected, a warning will be triggered or access will be blocked.

[0019] In this embodiment, the multi-source data of the hospital are uniformly integrated into the database, as follows: The hospital's multi-source data includes the electronic medical record system EMR, the laboratory information system LIS, the image archiving system PACS and the operation management system HIS; Among them, EMR stores patient electronic medical records, including text data (such as medical records) and structured data (such as test results); LIS stores laboratory test results, which are structured table data; PACS stores medical images, and the data is mainly in unstructured files (such as DICOM format), and also contains relevant metadata (such as patient information, image date); HIS stores hospital operation data, including financial systems and medical records, and the data is in the form of structured tables; The HL7 (Health Level Seven) protocol is used to standardize the interfaces of different data systems. Unstructured data (such as image files in PACS) are associated with structured data through metadata or indexing mechanisms; finally, the integrated data is stored in an integrated database.

[0020] In this embodiment, the sensitivity of the data in the database is graded by a data classification algorithm, and corresponding labels are automatically generated, as follows: Extract all tables, fields and record data from the database, including field names, field values, and table context information; Using regular matching, keyword recognition, and TF-IDF / NLP models, we calculate the comprehensive sensitivity score S of each field and table. The sensitivity levels are divided according to the thresholds, and labels are automatically generated and written into the meta-information table of the database.

[0021] In this embodiment, the comprehensive sensitivity score S is a weighted result of a plurality of characteristic indicators, as follows: Use predefined regular rules to match field names and field values ​​to determine whether they contain common personal privacy information. Use NLP models (such as Med-BERT or domain-trained BERT models) to perform semantic analysis on field names and context content to determine whether the field semantically belongs to personal identity information (PII). Based on the results of rule judgment and semantic analysis, assign a score S PII ; ; in, is the embedding vector of the field name; is a predefined PII classification template vector; It is the result of field name rule detection; is the result of semantic analysis; Analyze the content characteristics of fields or records, identify whether the field value contains privacy information or sensitive keywords, and calculate the content sensitivity score S content : ; Among them, W k is the keyword weight; Count(k,C i ) is the field value C i Total_Words(C i ) is the field value C i The total number of words; K is the total number of keywords; Combine the field name and the context semantics of the table to determine the sensitivity of the field content and obtain the context sensitivity score S context ; (semantic analysis results of field names and table structures on the core functions of fields); ; in, is the business sensitive weight; is the similarity between the field name and the sensitive semantic template; Determine the uniqueness and sensitivity of the field value based on its distribution and obtain the uniqueness score S uniqueness ; (For example, data with unique value distribution such as patient ID number and medical consultation number will get high scores); ; Among them, Unique_Count(F i ) is F i Total_Records(F i ) is the total number of records; Based on SPII , S content , S context and S uniqueness Calculate the comprehensive sensitivity score S total : ; in, is the weight coefficient.

[0022] In this embodiment, the desensitization model models the process of desensitizing user requests and data sensitive information as a multi-dimensional optimization problem. The goal is to formulate a global optimal desensitization strategy combination through dynamic analysis of user context, data features (field sensitivity) and access context (operation purpose or environment). At the same time, the model outputs the optimal desensitization rules and desensitized data, as follows: Define each F i , calculate F i The corresponding sensitivity score S i , user role R and the corresponding permission level Access_Threshold(R) (e.g. doctors can only see some desensitized fields, ADMIN can see all original data); access purpose M is the context (e.g. DIAGNOSIS, RESEARCH, statistical analysis, etc.); context adjustment factor C for sensitivity score m , used to dynamically relax or tighten sensitivity constraints; Each desensitization method Tj is represented by a set of parameters: Desensitization types: masking, generalization, encryption, replacement, perturbation; Desensitization degree: Strong desensitization L high , Medium Desensitization L medium , weak desensitization L low ; Each data field F i The selectable desensitization strategy set is T i ={T i1 ,T i2 ,...,T ij, ...,T iJ}; T ij, is the jth desensitization method; J is the total number of desensitization methods; The model goal is to select the embedding vector F of the field name i Desensitization strategy , the final optimization goal is: ; ; in, Score the security of the policy in the current context; The ability to preserve data availability for policy; The desensitization cost of the strategy (such as strategy complexity and computational cost); is the weight coefficient; To optimize the objective function.

[0023] In this embodiment, the desensitization model uses reinforcement learning DQL to continuously optimize the desensitization rules and improve the strategy based on user behavior data, as follows: The current state of the system S={S user ,S data ,S context} represents the context of the desensitization question, where S user is the user context, such as user role R, permission level, and user access behavior pattern; S data is the data context, for example, data field F i The sensitivity score S i , field distribution characteristics (uniqueness, distribution pattern); S context The access environment context includes the access purpose M (diagnosis and treatment, scientific research, etc.) and external environmental variables; The set of desensitization strategies selected for each data field is A. Each action corresponds to a desensitization method. Reinforcement learning DQL will select an optimal action for each field. By balancing data security and data availability, rewards That is, the optimization goal of the desensitization model: ; Among them, R safety Indicates whether data security requirements are met (whether sensitive fields are correctly desensitized). If the field sensitivity score S i If it is greater than the user permission threshold and is not desensitized, then ; If desensitization is correct, then R safety =1; R utility Whether the desensitized data still has statistical or business value is defined as the data availability score after desensitization; R complexity represents the cost of executing different desensitization strategies; is the weight coefficient; The strategy π represents the probability distribution of the action selected by the system each time: π(A|S;θ); Where θ is the parameter fitted by the deep neural network; Reinforcement learning DQL guides policy optimization through the value function Q(S,A): ; Among them, S′ is the new state reached after the current action; A′ is the set of actions corresponding to the new state; is the reward discount factor; Use a neural network to fit the Q function. Input the state S and output the Q(S,A) value for each action.

[0024] In this embodiment, the neural network training process is as follows: Record the historical state, action, reward, and state transition ; Randomly sample samples from the experience replay pool to train the network and reduce sample correlation; The training objective is to minimize the error between the value function and the current estimate of the DQN: ; ; Among them, are the target network parameters; represents the expected value of the sample (S,A,R,S′) in the experience replay sampling data; Y represents the target Q value; represents that the target network parameter is the Q value; represents that the target network parameter is the Q value; is the target function; By training the DQN, find the optimal action for each state S : .

[0025] In this embodiment, S4 is specifically: Input the user role R, access data field F i 、adjustment factor C m , and adjust the sensitivity score through the adjustment factor C m : ; Among them, S i ′ is the adjusted sensitivity score; Check whether the action meets the permissions. If not, return no access permission; if it meets, perform the desensitization policy utility calculation: Traverse T i ={T i1 ,T i2 ,...,T ij, ...,T iJ}, calculate the utility of each policy: ; For the utility score U(T ij ,F i ) takes the maximum value and selects the optimal strategy: ; Apply the optimal strategy to output the desensitized data.

[0026] In this embodiment, anomaly detection is performed based on the GANomaly model, based on the structure of the Generative Adversarial Network (GAN), and the reconstruction error, latent space consistency, and the discriminant ability of the discriminator are introduced to calculate the anomaly score, as follows: The encoder E maps the original input behavior feature X to a latent space vector z: ; in, The trainable parameters of the encoder; z is a latent space vector, which is a low-dimensional compressed representation of the behavior feature X; The generator G reconstructs the approximate features based on the latent space vector z generated by the encoder : ; in, is the trainable parameter of the generator; The discriminator D is responsible for distinguishing the input behavior features X from the features generated by the generator. , output a probability value D(X)∈[0,1] to determine whether the input is real data or generated reconstructed data:

[0027] in, is the activation function, is the weight matrix of the discriminator, is bias; In GANomaly, the loss function Loss due to reconstruction error , latent space consistency loss and the discriminative loss composition: ; ; ; ; Among them, λ 1 ,λ 2 is the weight coefficient, Features generated for the generator The output after inputting the encoder, Features generated for the generator The output after inputting the discriminator, is the expected function.

[0028] A hospital database management system based on adaptive data desensitization includes a processor, a memory, and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the hospital database management method based on adaptive data desensitization as described above.

[0029] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0030] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0031] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0032] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.

[0033] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. A hospital database management method based on adaptive data desensitization, characterized in that: The following steps are involved: S1: Integrate the hospital's multi-source data into a unified database; S2: Classify the sensitivity of the multi-source data in the database through a data classification algorithm and automatically generate corresponding labels; S3: Adjust the desensitization strategy in real time through context analysis and permission matching, establish a desensitization model based on role permissions, access content and data features, take user requests as input, and output the optimal combination of data desensitization rules; S4: The user queries data through the front-end system, enters relevant parameters, and determines whether the fields or data in the request contain highly sensitive fields. If they do, a suitable desensitization strategy is dynamically generated based on the desensitization model according to the user role, data sensitivity, and access context; S5: Combine data usage logs to monitor access behavior and perform anomaly detection based on the GANomaly model. If abnormal behavior is detected, a warning will be triggered or access will be blocked.

2. A hospital database management method based on adaptive data desensitization according to claim 1, characterized in that: The hospital's multi-source data is integrated into the database as follows: The hospital's multi-source data includes the electronic medical record system EMR, the laboratory information system LIS, the image archiving system PACS and the operation management system HIS; Among them, EMR stores patient electronic medical records, including text data and structured data; LIS stores laboratory test results, which are structured table data; PACS stores medical images, and the data is mainly unstructured files, and also contains relevant metadata; HIS stores hospital operation data, including financial systems and medical records, and the data is in the form of structured tables; The HL7 protocol is used to standardize the interfaces of different data systems, and unstructured data is associated with structured data through metadata or indexing mechanisms; finally, the integrated data is stored in an integrated database.

3. A hospital database management method based on adaptive data desensitization according to claim 1, characterized in that: The data classification algorithm is used to classify the sensitivity of the data in the database and automatically generate corresponding labels, as follows: Extract all tables, fields and record data from the database, including field names, field values, and table context information; Using regular matching, keyword recognition, and TF-IDF / NLP models, we calculate the comprehensive sensitivity score S of each field and table. The sensitivity levels are divided according to the thresholds, and labels are automatically generated and written into the meta-information table of the database.

4. A hospital database management method based on adaptive data desensitization according to claim 3, characterized in that: The comprehensive sensitivity score S is the weighted result of multiple characteristic indicators, as follows: Use predefined regular rules to match field names and field values ​​to determine whether they contain common personal privacy information. Use the NLP model to perform semantic analysis on field names and context content to determine whether the field semantically belongs to personal identity information (PII). Based on the results of rule judgment and semantic analysis, assign a score S. PII ; ; in, is the embedding vector of the field name; is a predefined PII classification template vector; It is the result of field name rule detection; is the result of semantic analysis; Analyze the content characteristics of fields or records, identify whether the field value contains privacy information or sensitive keywords, and calculate the content sensitivity score S content : ; Among them, W k is the keyword weight; Count(k,C i ) is the field value C i Total_Words(C i ) is the field value C i The total number of words; K is the total number of keywords; Combine the field name and the context semantics of the table to determine the sensitivity of the field content and obtain the context sensitivity score S context ; ; in, is the business sensitive weight; is the similarity between the field name and the sensitive semantic template; Determine the uniqueness and sensitivity of the field value based on its distribution and obtain the uniqueness score S uniqueness ; ; Among them, Unique_Count(F i ) is F i Total_Records(F i ) is the total number of records; Based on S PII , S content , S context and S uniqueness Calculate the comprehensive sensitivity score S total : ; in, is the weight coefficient.

5. The hospital database management method based on adaptive data desensitization according to claim 1 is characterized in that: The desensitization model models the process of desensitizing user requests and data sensitive information as a multi-dimensional optimization problem. The goal is to formulate a global optimal desensitization strategy combination through dynamic analysis of user context, data features, and access context. At the same time, the model outputs the optimal desensitization rules and desensitized data, as follows: Define each F i , calculate F i The corresponding sensitivity score S i , user role R and the corresponding permission level Access_Threshold(R); access purpose M is the context; context adjustment factor C for sensitivity score m , used to dynamically relax or tighten sensitivity constraints; The model goal is to select the embedding vector F of the field name i Desensitization strategy , the final optimization goal is: ; ; in, Score the security of the policy in the current context; The ability to preserve data availability for policy; The desensitization cost of the strategy; is the weight coefficient; To optimize the objective function.

6. A hospital database management method based on adaptive data desensitization according to claim 5, characterized in that: The desensitization model uses reinforcement learning DQL to continuously optimize desensitization rules and improve strategies based on user behavior data, as follows: The current state of the system S={S user ,S data ,S context } represents the context of the desensitization question, where S user is the user context; S data is the data context; S context To access the environment context; The set of desensitization strategies selected for each data field is A. Each action corresponds to a desensitization method. Reinforcement learning DQL will select an optimal action for each field. By balancing data security and data availability, rewards That is, the optimization goal of the desensitization model: ; Among them, R safety Indicates whether the data security requirements are met. If the field sensitivity score S i If it is greater than the user permission threshold and is not desensitized, ; If desensitization is correct, then R safety =1; R utility Whether the desensitized data still has statistical or business value is defined as the data availability score after desensitization; R complexity represents the cost of executing different desensitization strategies; is the weight coefficient; The strategy π represents the probability distribution of the action selected by the system each time: π(A|S;θ); Where θ is the parameter fitted by the deep neural network; Reinforcement learning DQL guides policy optimization through the value function Q(S,A): ; Among them, S' is the new state reached after the current action; A' is the action set corresponding to the new state; is the reward discount factor; Use a neural network to fit the Q function, input the state S, and output the Q(S,A) value of each action.

7. A hospital database management method based on adaptive data desensitization according to claim 6, characterized in that: The training process of the neural network is as follows: Record historical status, actions, rewards, and state transfers ; Randomly extract samples from the experience replay pool to train the network and reduce sample correlation; The training target is the error between the value function and the current estimate of DQN: ; ; in, is the target network parameter; represents the expected value of the sample (S, A, R, S′) in the experience playback sampling data; Y represents the target Q value; Indicates that the target network parameters are Q value; Indicates that the target network parameters are Q value; is the objective function; By training DQN, the optimal action is found in each state S : 。 8. The hospital database management method based on adaptive data desensitization according to claim 5 is characterized in that: The S4 is specifically: Enter user role R, access data field F i , adjustment factor C m , and by adjusting the factor C m Adjust the sensitivity score: ; Among them, S i ′Adjusted sensitivity score; Check Action Whether the permission is satisfied, if not, it returns no access permission; if satisfied, the desensitization strategy utility calculation is performed: Traverse T i ={T i1 ,T i2 ,...,T ij, ...,T iJ }, calculate the utility of each strategy: ; For the utility score U(T ij ,F i ) takes the maximum value and selects the optimal strategy: ; Apply the optimal strategy to output the desensitized data.

9. The hospital database management method based on adaptive data desensitization according to claim 1 is characterized in that: The anomaly detection based on the GANomaly model is based on the structure of the generative adversarial network GAN, and introduces reconstruction error, latent space consistency, and the discriminant ability of the discriminator to calculate the anomaly score, as follows: The encoder E maps the original input behavior feature X to a latent space vector z: ; in, The trainable parameters of the encoder; z is a latent space vector, which is a low-dimensional compressed representation of the behavior feature X; The generator G reconstructs the approximate features based on the latent space vector z generated by the encoder : ; in, is the trainable parameter of the generator; The discriminator D is responsible for distinguishing the input behavior features X from the features generated by the generator. , output a probability value D(X)∈[0,1] to determine whether the input is real data or generated reconstructed data: ; in, is the activation function, is the weight matrix of the discriminator, is bias; In GANomaly, the loss function Loss due to reconstruction error , latent space consistency loss and the discriminative loss composition: ; ; ; ; Among them, λ1,λ2 are weight coefficients, Features generated for the generator The output after inputting the encoder, Features generated for the generator The output after inputting the discriminator, is the expected function.

10. A hospital database management system based on adaptive data desensitization, characterized in that: It includes a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the hospital database management method based on adaptive data desensitization as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data desensitization method and system

    CN115391814A

  • Interface data desensitization processing method and device

    CN117592108A

  • Dynamic data adaptive desensitization method and device based on artificial intelligence

    CN119128990A

  • Learning to Transform Sensitive Data with Variable Distribution Preservation

    US20220311749A1

Cited By

  • Method and platform for integrating and sharing health data of old people

    CN120321057A

  • Business candidate person and member information management and storage method and system

    CN120372669A

  • Data security management method and system

    CN120434011A

  • Real-time data desensitization method and system based on context awareness

    CN120750643A

  • A context-aware real-time data anonymization method and system

    CN120750643B