A Hospital Database Management Method and System Based on Adaptive Data Masking
Through the adaptive data desensitization method, combined with sensitivity grading, desensitization model and abnormal detection, the problem of data security and availability imbalance in traditional methods is solved, and efficient utilization and intelligent management of hospital data are achieved.
Patent Information
- Application Number
- CN202510511995.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Traditional static data desensitization methods are difficult to adapt to complex hospital scenarios, resulting in an imbalance between data security and availability, and cannot meet the needs of data protection and diversified use.
Adaptive data desensitization method is adopted, sensitivity hierarchy is performed through data classification algorithm, and a desensitization model based on role permissions, access content and data characteristics is established. Desensitization rules are optimized in combination with reinforcement learning, and abnormal detection is used using the GANomaly model to dynamically generate the optimal desensitization strategy.
It realizes efficient use and sharing of hospital data while ensuring data security, improves the intelligence level of data management, dynamically adjusts the desensitization strategy to adapt to different user roles and access scenarios, and accurately detects abnormal behaviors.
Smart Images

Figure CN120030601B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management, and particularly to a hospital database management method and system based on adaptive data masking. Background Art
[0002] With the rapid development of medical informatization, a large amount of sensitive data containing patient privacy has been accumulated in the daily operation of hospitals. These data mainly come from the Electronic Medical Record System (EMR), Laboratory Information System (LIS), Picture Archiving and Communication System (PACS), and Hospital Information System (HIS). These data provide support for hospital diagnosis and treatment, scientific research, operation, etc. However, their highly sensitive attributes (such as patient names, ID numbers, medical records, etc.) make them face high risks of data leakage, abuse, and illegal access. At the same time, in the context of big data analysis, artificial intelligence applications, and test environments, the open demand for data further increases the possibility of data leakage.
[0003] With the increasing attention to data sensitivity, higher requirements are put forward for data management: not only need to ensure that sensitive data is not leaked, but also need to meet diverse usage requirements on the basis of compliance. However, traditional static data masking methods are difficult to adapt to the complex scenarios of hospitals and are prone to imbalance between security and usability. Summary of the Invention
[0004] In order to solve the above problems, the purpose of the present invention is to provide a hospital database management method and system based on adaptive data masking, which can not only meet the hospital's requirements in data protection and privacy compliance, but also realize the efficient utilization and sharing of data on the premise of ensuring data security, and improve the intelligent level of hospital data management.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A hospital database management method based on adaptive data masking, comprising the following steps:
[0007] S1: Integrate the multi-source data of the hospital into the database uniformly;
[0008] S2: Classify the sensitivity of the multi-source data in the database through a data classification algorithm and automatically generate corresponding labels;
[0009] S3: Adjust the masking strategy in real time through context analysis and permission matching, establish a masking model based on role permissions, access content, and data characteristics, take the user request as input, and output the optimal data masking rule combination;
[0010] S4: The user queries data through the front-end system, enters relevant parameters, and determines whether the fields or data in the request contain highly sensitive fields. If highly sensitive fields are included, a suitable desensitization strategy is dynamically generated based on the user role, data sensitivity, and access context, using a desensitization model.
[0011] S5: Combine data usage logs to monitor access behavior, and perform anomaly detection based on the GANomaly model. If abnormal behavior is detected, a warning is triggered or access is blocked.
[0012] Furthermore, the multi-source data of the hospital is unified and integrated into the database, as follows:
[0013] The multi-source data of the hospital includes the Electronic Medical Record system (EMR), Laboratory Information System (LIS), Picture Archiving and Communication System (PACS), and Hospital Information System (HIS).
[0014] Among them, the EMR stores patients' electronic medical records, including text data and structured data; the LIS stores laboratory test results, which are structured tabular data; the PACS stores medical images, with the data mainly in unstructured files and including relevant metadata; the HIS stores hospital operation data, including the financial system and medical records, and the data form is structured tables.
[0015] Use the HL7 protocol to standardize the interfaces of different data systems, and associate unstructured data with structured data through metadata or indexing mechanisms; finally, store the integrated data in an integrated database.
[0016] Furthermore, the data in the database is classified by a data classification algorithm according to sensitivity levels, and corresponding labels are automatically generated, as follows:
[0017] Extract all tables, fields, and record data from the database, including field names, field values, and table context information.
[0018] Use regular expression matching, keyword recognition, and TF-IDF / NLP models to calculate the comprehensive sensitivity score S for each field and table.
[0019] Divide the sensitivity levels according to the threshold, automatically generate labels, and write them into the metadata table of the database.
[0020] Furthermore, the comprehensive sensitivity score S is the weighted result of multiple characteristic indicators, as follows:
[0021] Match the field name and field value using predefined regular rules to determine whether they contain common personal privacy information. Use an NLP model to perform semantic analysis on the field name and context content to determine whether the field semantically belongs to personal identity information (PII). Assign a score S based on the rule judgment and semantic analysis results PII ;
[0022] ;
[0023] Among them, is the embedding vector of the field name; is the predefined PII classification template vector; is the result of field name rule detection; is the semantic analysis result;
[0024] Analyze the content characteristics of the field or record, identify whether the field value contains privacy information or sensitive keywords, and calculate the content sensitivity score S content :
[0025] ;
[0026] Among them, W k is the keyword weight; Count(k, C i ) is the number of occurrences of keyword k in the field value C i ; Total_Words(C i ) is the total number of words in the field value C i ; K is the total number of keywords;
[0027] Combine the field name and the context semantics of the table to judge the sensitivity of the field content and obtain the context sensitivity score S context ;
[0028] ;
[0029] Among them, is the business sensitivity weight; is the similarity between the field name and the sensitive semantic template;
[0030] Judge the uniqueness and sensitivity according to the distribution of the field value, and obtain the uniqueness score S uniqueness ;
[0031] ;
[0032] Among them, Unique_Count(F i ) is the number of unique values of F i ; Total_Records(F i ) is the total number of records;
[0033] Based on S PII , S content , S context and S uniqueness Calculate the comprehensive sensitivity score S total :
[0034] ;
[0035] where α, β, γ, δ are weight coefficients.
[0036] Furthermore, the desensitization model models the desensitization process of user requests and data sensitive information as a multi-dimensional optimization problem. The goal is to formulate a globally optimal desensitization strategy combination through the dynamic analysis of user context, data characteristics, and access context. At the same time, the model outputs the optimal desensitization rules and the desensitized data, as follows:
[0037] Define each F i , calculate the sensitivity score S i corresponding to F i , the user role R and the permission level Access_Threshold(R) corresponding to the user role; the access purpose M is the context environment; the adjustment factor C m of the context for the sensitivity score, which is used to dynamically relax or tighten the sensitivity limit;
[0038] The model goal is to select the desensitization strategy i of the embedding vector F of the field name, and the final optimization goal is:
[0039] ;
[0040] ;
[0041] where is the security score of the strategy in the current context; is the ability of the strategy to retain data availability; is the desensitization cost of the strategy; is the weight coefficient; is the optimization objective function.
[0042] Furthermore, the desensitization model uses deep Q-learning (DQL) to continuously optimize the desensitization rules and improve the strategy based on user behavior data, as follows:
[0043] The current state of the system S = {S user , S data , S context} represents the context of the desensitization problem, where S useris the user context; S data is the data context; S context is the access environment context;
[0044] The set of desensitization strategies selected for each data field is A, and each action corresponds to a desensitization method. The deep Q - learning (DQL) will select an optimal action for each field;
[0045] By balancing data security and data availability, the reward is the optimization goal of the desensitization model:
[0046] ;
[0047] Among them, R safety represents whether the data security requirements are met. If the field sensitivity score S i is greater than the user permission threshold and not desensitized, then ; if the desensitization is correct, then R safety = 1; R utility is whether the desensitized data still has statistical or business value, defined as the data availability score of the desensitized data; represents the cost of executing different desensitization strategies; is the weight coefficient;
[0048] The policy π represents the probability distribution of the system's action selection each time:
[0049] π(A∣S;θ);
[0050] where θ is the parameter fitted by the deep neural network;
[0051] The deep Q - learning (DQL) guides policy optimization through the value function Q(S,A):
[0052] ;
[0053] Among them, S′ is the new state reached after the current action; A′ is the set of actions corresponding to the new state; is the reward discount factor;
[0054] Use the neural network to fit the Q - function, input the state S, and output the Q(S,A) value of each action.
[0055] Furthermore, the training process of the neural network is as follows:
[0056] Record the historical states, actions, rewards, and state transitions ; randomly sample from the experience replay pool to train the network and reduce sample correlation;
[0057] The training objective is to minimize the error between the value function and the current estimate of DQN:
[0058] ;
[0059] ;
[0060] where, are the target network parameters; represents the expected value of the sample (S, A, R, S′) in the sampled data of the experience replay; Y represents the target Q value; represents that the target network parameter is of the Q value; represents that the target network parameter is of the Q value; is the target function;
[0061] By training DQN, the optimal action A can be found for each state S ∗ :
[0062] .
[0063] Furthermore, S4 is specifically:
[0064] Input the user role R, the access data field F i , the adjustment factor C m , and adjust the sensitivity score through the adjustment factor C m :
[0065] ;
[0066] where, S i ′ is the adjusted sensitivity score;
[0067] Check whether the action meets the permissions. If not, return no access permission; if it meets, perform the desensitization policy utility calculation:
[0068] Traverse T i ={T i1 , T i2 ,..., T ij, ..., T iJ}, and calculate the utility of each policy:
[0069] ;
[0070] Take the maximum value of the utility score U(T ij , F i ) to select the optimal policy:
[0071] ;
[0072] Output the desensitized data by applying the optimal strategy.
[0073] Furthermore, perform anomaly detection based on the GANomaly model. Based on the structure of the generative adversarial network GAN, introduce the reconstruction error, latent space consistency, and discriminative ability of the discriminator to calculate the anomaly score, as follows:
[0074] The encoder E maps the original input behavior feature X to the latent space vector z:
[0075] ;
[0076] where are the trainable parameters of the encoder; z is the latent space vector, which is the low-dimensional compressed representation of the behavior feature X;
[0077] The generator G reconstructs an approximate feature according to the latent space vector z generated by the encoder :
[0078] ;
[0079] where are the trainable parameters of the generator;
[0080] The discriminator D is responsible for distinguishing between the input behavior feature X and the feature generated by the generator , and outputs a probability value D(X) ∈ [0, 1] to determine whether the input is real data or the generated reconstructed data:
[0081]
[0082] where is the activation function, is the weight matrix of the discriminator, is the bias;
[0083] In GANomaly, the loss function consists of the reconstruction error loss , the latent space consistency loss and the discriminative loss :
[0084] ;
[0085] ;
[0086] ;
[0087] ;
[0088] where λ1 and λ2 are weight coefficients, is the feature generated by the generator which is the output after inputting into the encoder, is the feature generated by the generator which is the output after inputting into the discriminator, is the expectation function.
[0089] A hospital database management system based on adaptive data desensitization, comprising a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned hospital database management method based on adaptive data desensitization.
[0090] The present invention has the following beneficial effects:
[0091] 1. The present invention integrates PII recognition, content sensitivity analysis, context semantic analysis, and uniqueness detection through a sensitivity grading model, automatically divides the hospital database data into highly sensitive, medium sensitive, and low sensitive levels, and simultaneously generates traceable sensitivity labels, providing a technical basis for data security, dynamic desensitization, and compliance management;
[0092] 2. The present invention can dynamically generate an optimal desensitization strategy according to the user role, access content, and scenarios (such as diagnosis and treatment, scientific research, testing), avoiding the "one-size-fits-all" problem of static desensitization rules; and continuously optimizes the desensitization rules through reinforcement learning (such as Deep Q-Learning), can continuously learn and improve the strategy according to the user behavior data, and make the desensitization effect more accurate;
[0093] 3. The present invention is based on the GANomaly model, through the collaborative work of the encoder, generator, and discriminator, uses normal behavior modeling as a baseline to capture abnormal behaviors deviating from the normal mode, and through the reconstruction error and latent space consistency, can accurately detect abnormal data queries, unauthorized behaviors, or batch sensitive data accesses, thereby protecting data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0095] The following further describes the present invention in detail with reference to the accompanying drawings and specific embodiments:
[0096] Refer to Figure 1 , in this embodiment, a hospital database management method based on adaptive data desensitization is provided, including the following steps:
[0097] S1: Integrate the multi-source data of the hospital into the database uniformly;
[0098] S2: Classify the sensitivity of the multi-source data in the database through a data classification algorithm, and automatically generate corresponding tags;
[0099] S3: Adjust the data masking strategy in real time through context analysis and permission matching, establish a data masking model based on role permissions, access content, and data characteristics, take the user request as input, and output the optimal combination of data masking rules;
[0100] S4: The user queries data through the front-end system, inputs relevant parameters, and determines whether the fields or data in the request contain highly sensitive fields. If highly sensitive fields are included, based on the user role, data sensitivity, and access context, a suitable data masking strategy is dynamically generated based on the data masking model;
[0101] S5: Combine the data usage logs to monitor access behaviors, perform anomaly detection based on the GANomaly model. If abnormal behaviors (such as unauthorized queries) are detected, trigger warnings or block access.
[0102] In this embodiment, integrating the multi-source data of the hospital into the database uniformly is as follows:
[0103] The multi-source data of the hospital includes the Electronic Medical Record System (EMR), Laboratory Information System (LIS), Picture Archiving and Communication System (PACS), and Hospital Information System (HIS);
[0104] Among them, EMR stores patients' electronic medical records, including text data (such as medical record records) and structured data (such as examination results); LIS stores laboratory examination results, which are structured tabular data; PACS stores medical images, and the data is mainly in the form of unstructured files (such as DICOM format), and also includes relevant metadata (such as patient information, imaging date); HIS stores hospital operation data, including the financial system and medical records, and the data form is structured tables;
[0105] Use the HL7 (Health Level Seven) protocol to standardize the interfaces of different data systems. Unstructured data (such as image files in PACS) is associated with structured data through metadata or indexing mechanisms; finally, the integrated data is stored in an integrated database.
[0106] In this embodiment, classifying the sensitivity of the data in the database through a data classification algorithm and automatically generating corresponding tags is as follows:
[0107] Extract all tables, fields, and record data from the database, including field names, field values, and table context information;
[0108] Using regular matching, keyword recognition, and TF-IDF / NLP models, calculate the comprehensive sensitivity score S for each field and table;
[0109] Divide the sensitivity levels by a threshold, automatically generate labels, and write them into the meta-information table of the database.
[0110] In this embodiment, the comprehensive sensitivity score S is the weighted result of integrating multiple feature indicators, specifically as follows:
[0111] Use predefined regular rules to match field names and field values to determine whether they contain common personal privacy information. Use an NLP model (such as Med-BERT or a domain self-trained BERT model) to perform semantic analysis on the field name and context content to determine whether the field semantically belongs to personal identity information PII. Assign a score S based on the rule judgment and semantic analysis results PII ;
[0112] ;
[0113] where, is the embedding vector of the field name; is the predefined PII classification template vector; is the field name rule detection result; is the semantic analysis result;
[0114] Analyze the content features of the field or record, identify whether the field value contains privacy information or sensitive keywords, and calculate the content sensitivity score S content :
[0115] ;
[0116] where, W k is the keyword weight; Count(k, C i ) is the number of occurrences of keyword k in field value C i ; Total_Words(C i ) is the total number of words in field value C i ; K is the total number of keywords;
[0117] Combine the context semantics of the field name and the table to judge the sensitivity of the field content and obtain the context sensitivity score S context ; (the semantic analysis result of the field name and table structure on the core function of the field);
[0118] ;
[0119] where, is the business sensitivity weight; The similarity between the field name and the sensitive semantic template;
[0120] Judge its uniqueness and sensitivity according to the distribution of the field value, and obtain the uniqueness score S uniqueness ; (Data with a unique value distribution such as patient ID number, visit number, etc. will get a high score);
[0121] ;
[0122] Among them, Unique_Count(F i ) is the number of unique values of F i ; Total_Records(F i ) is the total number of records;
[0123] Based on S PII , S content , S context and S uniqueness calculate the comprehensive sensitivity score S total :
[0124] ;
[0125] Among them, α, β, γ, δ are weight coefficients.
[0126] In this embodiment, the desensitization model models the user request and the data sensitive information desensitization process as a multi-dimensional optimization problem. The goal is to formulate a globally optimal desensitization strategy combination through the dynamic analysis of user context, data characteristics (field sensitivity), and access context (operation purpose or environment). At the same time, the model outputs the optimal desensitization rules and the desensitized data, as follows:
[0127] Define each F i , calculate the corresponding sensitivity score S i of F i , user role R and the permission level Access_Threshold(R) corresponding to the user role (for example, doctors can only view some desensitized fields, and ADMIN can view all original data); access purpose M is the context environment (such as DIAGNOSIS medical access, RESEARCH scientific research access, statistical analysis, etc.); context adjustment factor C m for dynamically relaxing or tightening sensitivity restrictions;
[0128] Each desensitization method Tj is represented by a set of parameters:
[0129] Desensitization type: masking, generalization, encryption, replacement, perturbation;
[0130] Degree of desensitization: strong desensitization Lhigh , Medium desensitization L medium , Weak desensitization L low ;
[0131] Each data field F i The set of optional desensitization strategies is T i ={T i1 , T i2 ,..., T ij, ..., T iJ}; T ij, is the j-th desensitization method; J is the total number of desensitization methods;
[0132] The model goal is to select the desensitization strategy of the embedding vector F of the field name i , and the final optimization goal is:
[0133] ;
[0134] ;
[0135] Among them, is the security score of the strategy in the current context; is the ability of the strategy to retain data availability; is the desensitization cost of the strategy (such as strategy complexity and computational cost); is the weight coefficient; is the optimization objective function.
[0136] In this embodiment, the desensitization model uses deep Q-learning (DQL) of reinforcement learning to continuously optimize the desensitization rules and improve the strategy based on user behavior data, specifically as follows:
[0137] The current state of the system S = {S user , S data , S context} represents the context of the desensitization problem, where S user is the user context, such as user role R, permission level, user access behavior pattern; S data is the data context, such as the sensitivity score S i of the data field F i , field distribution characteristics (uniqueness, distribution pattern); S context is the access environment context, including the access purpose M (diagnosis and treatment, scientific research, etc.) and external environmental variables;
[0138] The set of desensitization strategies selected for each data field is A, and each action corresponds to a desensitization method. The deep Q-learning (DQL) of reinforcement learning will select an optimal action for each field;
[0139] By balancing data security and data availability, rewards That is, the optimization goal of the desensitization model:
[0140] ;
[0141] Among them, R safety Indicates whether the data security requirements are met (whether sensitive fields are correctly desensitized). If the field sensitivity score S i Is greater than the user permission threshold and not desensitized, then ; If the desensitization is correct, then R safety = 1; R utility Is whether the desensitized data still has statistical or business value, defined as the data availability score after desensitization; Represents the cost of executing different desensitization strategies; Is the weight coefficient;
[0142] The policy π represents the probability distribution of the system's action selection each time:
[0143] π(A∣S;θ);
[0144] Among them, θ is the parameter fitted by the deep neural network;
[0145] The reinforcement learning DQL guides policy optimization through the value function Q(S,A):
[0146] ;
[0147] Among them, S′ is the new state reached after the current action; A′ is the set of actions corresponding to the new state; Is the reward discount factor;
[0148] Use the neural network to fit the Q function, input the state S, and output the Q(S,A) value of each action.
[0149] In this embodiment, the neural network training process is as follows:
[0150] Record historical states, actions, rewards, and state transitions ; Randomly sample from the experience replay pool to train the network and reduce sample correlation;
[0151] The training goal is to minimize the error between the value function and the current estimate of the DQN:
[0152] ;
[0153] ;
[0154] Among them, Is the target network parameter; denotes the expected value of the sample (S, A, R, S′) in the experience replay sampling data; Y denotes the target Q value; denotes that the target network parameter is the Q value of; denotes that the target network parameter is the Q value of; is the target function;
[0155] By training the DQN, the optimal action A can be found for each state S ∗ :
[0156] .
[0157] In this embodiment, S4 is specifically:
[0158] Input the user role R, access data field F i , adjustment factor C m , and adjust the sensitivity score through the adjustment factor C m :
[0159] ;
[0160] where S i ′ is the adjusted sensitivity score;
[0161] Check the action to see if it meets the permissions. If not, return no access permission; if so, perform the desensitization policy utility calculation:
[0162] Traverse T i ={T i1 , T i2 ,..., T ij, ..., T iJ}, and calculate the utility of each policy:
[0163] ;
[0164] Take the maximum value of the utility score U(T ij , F i ) to select the optimal policy:
[0165] ;
[0166] Apply the optimal policy to output the desensitized data.
[0167] In this embodiment, anomaly detection is performed based on the GANomaly model. Based on the structure of the generative adversarial network (GAN), the reconstruction error, the latent space consistency, and the discriminative ability of the discriminator are introduced to calculate the anomaly score, as follows:
[0168] The encoder E maps the original input behavior feature X to the latent space vector z:
[0169] ;
[0170] where, are the trainable parameters of the encoder; z is the latent space vector, which is the low-dimensional compressed representation of the behavior feature X;
[0171] The generator G reconstructs an approximate feature according to the latent space vector z generated by the encoder :
[0172] ;
[0173] where, are the trainable parameters of the generator;
[0174] The discriminator D is responsible for distinguishing the input behavior feature X and the feature generated by the generator , and outputs a probability value D(X) ∈ [0, 1] to determine whether the input is real data or the generated reconstructed data:
[0175]
[0176] where, is the activation function, is the weight matrix of the discriminator, is the bias;
[0177] In GANomaly, the loss function consists of the reconstruction error loss , the latent space consistency loss and the discriminative loss :
[0178] ;
[0179] ;
[0180] ;
[0181] ;
[0182] where, λ1, λ2 are the weight coefficients, is the feature generated by the generator The output after the input encoder, is the feature generated by the generator The output after the input discriminator, is the desired function.
[0183] A hospital database management system based on adaptive data desensitization includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned method for managing a hospital database based on adaptive data desensitization.
[0184] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0185] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0186] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide means for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1Steps of the functions specified in one or more boxes.
[0188] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation of the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A hospital database management method based on adaptive data desensitization, characterized in that, It includes the following steps: S1: Uniformly integrate the multi-source data of the hospital into the database; S2: For the multi-source data in the database, perform sensitivity grading on the database data through a data classification algorithm and automatically generate corresponding tags; S3: Real-time adjust the desensitization policy through context analysis and permission matching, establish a desensitization model based on role permissions, access content, and data characteristics, take the user request as input, and output the optimal data desensitization rule combination; S4: The user queries data through the front-end system, inputs relevant parameters, and determines whether the fields or data in the request contain highly sensitive fields. If highly sensitive fields are included, based on the user role, data sensitivity, and access context, dynamically generate an appropriate desensitization policy based on the desensitization model; S5: Combine the data usage log records to monitor access behaviors, perform anomaly detection based on the GANomaly model. If an abnormal behavior is detected, trigger a warning or block the access; The desensitization model models the user request and the data sensitive information desensitization process as a multi-dimensional optimization problem. The goal is to formulate a globally optimal desensitization policy combination through the dynamic analysis of user context, data characteristics, and access context. At the same time, the model outputs the optimal desensitization rules and the desensitized data, specifically as follows: Define the embedding vector F for each field name i , calculate F i The corresponding comprehensive sensitivity score S, user role RU, and the access level Access_Threshold(RU) corresponding to the user role; the adjustment factor C of the context for the sensitivity score m , which is used to dynamically relax or tighten the sensitivity limit; The model objective is to select F i for the desensitization strategy , and the final optimization objective is as follows: ; ; Among them, is the security score of the policy in the current context; is the retention ability of the policy for data availability; is the desensitization cost of the policy; is the weight coefficient; is the optimization objective function; The specific content of S4 is as follows: Input user role RU, embedding vector F of the access data field i , adjustment factor C m , and adjust the comprehensive sensitivity score through the adjustment factor C m : ; Among them, S′ is the adjusted comprehensive sensitivity score; Check action Check if the permissions are satisfied. If not, return "No access permission". If satisfied, calculate the utility of the desensitization policy: Traverse T i ={T i1 ,T i2 ,...,T ij, ...,T iJ}, calculate the utility of each strategy: ; Take the maximum value of the utility score U(T ij ,F i ) and select the optimal strategy: ; Apply the optimal policy to output the desensitized data.
2. The hospital database management method based on adaptive data desensitization according to claim 1, characterized in that, The specific content of uniformly integrating the multi-source data of the hospital into the database is as follows: The multi-source data of the hospital includes the Electronic Medical Record system (EMR), Laboratory Information System (LIS), Picture Archiving and Communication System (PACS), and Hospital Information System (HIS); Among them, EMR stores patients' electronic medical records, including text data and structured data; LIS stores laboratory test results, which are structured tabular data; PACS stores medical images, with unstructured files as the main body of the data, and also includes relevant metadata; HIS stores hospital operation data, including the financial system and medical records, and the data form is structured tables; Use the HL7 protocol to standardize the interfaces of different data systems, and associate unstructured data with structured data through metadata or indexing mechanisms; finally, store the integrated data in an integrated database.
3. A hospital database management method based on adaptive data desensitization according to claim 1, characterized in that, The specific content of performing sensitivity grading on the database data through a data classification algorithm and automatically generating corresponding tags is as follows: Extract all tables, fields, and record data from the database, including field names, field values, and table context information; Use regular matching, keyword recognition, and TF-IDF / NLP models to calculate the comprehensive sensitivity score S of each field and table; Divide the sensitivity levels according to the threshold, automatically generate tags and write them into the metadata table of the database.
4. A hospital database management method based on adaptive data desensitization according to claim 3, characterized in that, The comprehensive sensitivity score S is the weighted result of integrating multiple characteristic indicators, specifically as follows: Match the field name and field value using predefined regular rules to determine whether they contain common personal privacy information. Use an NLP model to perform semantic analysis on the field name and the context content to determine whether the field semantically belongs to personal identity information PII. Assign a score S based on the results obtained from the rule judgment and semantic analysis PII ; ; Among them, is the embedding vector of the field name; is the predefined PII classification template vector; is the detection result of the field name rule; is the semantic analysis result; Analyze the content characteristics of the field or record, identify whether the field value contains privacy information or sensitive keywords, and calculate the content sensitivity score S content : ; Among them, W k is the keyword weight; Count(k, C i ) is the number of occurrences of keyword k in field value C i ; Total_Words(C i ) is the total number of words in field value C i ; K is the total number of keywords; Based on the context semantics of the field name and the table, determine the sensitivity of the field content and obtain the context-sensitive score S context ; ; Among them, is the business-sensitive weight; is the similarity between the field name and the sensitive semantic template; Determine its uniqueness and sensitivity based on the distribution of field values, and obtain the uniqueness score S uniqueness ; ; Among them, Unique_Count(F i ) is the number of unique values of F i ; Total_Records(F i ) is the total number of records; Based on S PII 、S content 、S context and S uniqueness Calculate the comprehensive sensitivity score S: ; Among them, α, β, γ, δ are weight coefficients.
5. A hospital database management method based on adaptive data desensitization according to claim 1, characterized in that, The desensitization model uses Deep Q-Learning (DQL) of reinforcement learning to continuously optimize the desensitization rules and improve the policy based on user behavior data, specifically as follows: The current state S of the system u ={S user ,S data ,S text}, which represents the context of the de - sensitization problem, where S user is the user context; S data is the data context; S text is the access environment context; The set of desensitization strategies selected for each data field is A, and each action corresponds to a desensitization method. The deep Q-learning (DQL) will select an optimal action for each field; By balancing data security and data availability, reward That is, the optimization goal of the desensitization model: ; Among them, R safety indicates whether the data security requirements are met. If the comprehensive sensitivity score S is greater than the user permission threshold and not desensitized, then ; if the desensitization is correct, then R safety = 1; R utility represents whether the desensitized data still has statistical or business value, defined as the data availability score after desensitization; represents the cost of executing different desensitization strategies; is the weight coefficient; The policy π represents the probability distribution of the actions selected by the system each time: π(A∣S u ;θ); where θ is the parameter fitted by the deep neural network; Deep Q-Learning guides policy optimization through the value function Q(S u , A): ; Among them, S′ is the new state reached after the current action; A′ is the set of actions corresponding to the new state; is the reward discount factor; Use a deep neural network to fit the Q-function, with the input state S u , and output the Q(S u , A) value for each action.
6. The hospital database management method based on adaptive data desensitization according to claim 5, wherein The training process of the neural network is as follows: Record the state, action, reward, and state transition of history ; Randomly sample from the experience replay pool to train the network and reduce sample correlation; The training objective is the error between the value function and the current estimate of the DQN: ; ; Among them, is the target network parameter; represents the expected value of the sample (S u , A, R, S′) in the experience replay sampled data; Y represents the target Q value; represents that the target network parameter is the Q value of; represents that the target network parameter is the Q value of; is the objective function; By training the DQN, the optimal action can be found under state S u : 。 7. A hospital database management method based on adaptive data desensitization according to claim 1, characterized in that The anomaly detection based on the GANomaly model is carried out based on the structure of the generative adversarial network (GAN), and the reconstruction error, the consistency of the latent space, and the discriminative ability of the discriminator are introduced to calculate the anomaly score, which is as follows: The encoder E maps the original input behavior feature X to the latent space vector z: ; Among them, the trainable parameters of the encoder; z is the latent space vector, which is the low-dimensional compressed representation of the behavioral feature X; The generator G reconstructs approximate features based on the latent space vector z generated by the encoder. : ; Among them, are the trainable parameters of the generator; The discriminator D is responsible for distinguishing the input behavioral feature X from the feature generated by the generator, and outputs a probability value D(X) ∈ [0, 1] to determine whether the input is real data or the generated reconstructed data: ; Among them, is the activation function, is the weight matrix of the discriminator, is the bias; In GANomaly, the loss function consists of a reconstruction error loss , a latent space consistency loss and a discriminator loss : ; ; ; ; where λ1 and λ2 are weight coefficients, are the features generated by the generator which are the outputs after being input into the encoder, are the features generated by the generator which are the outputs after being input into the discriminator, is the expected function.
8. A hospital database management system based on adaptive data desensitization, characterized in that, It includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in a hospital database management method based on adaptive data desensitization according to any one of claims 1-7.
Citation Information
Patent Citations
Data desensitization method and system
CN115391814A
Dynamic data adaptive desensitization method and device based on artificial intelligence
CN119128990A