Abnormal diagnosis and treatment behavior detection method and device, server and computer storage medium
Through federated learning and hash alignment technology, combined with model parameters cutting and noise processing, the problem of abnormal diagnosis and treatment behavior detection in the medical industry is solved, and effective protection of patient privacy and efficient detection of diagnosis and treatment behavior is achieved.
Patent Information
- Application Number
- CN202510154986.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively detect and prevent abnormal diagnosis and treatment behaviors in the medical industry, especially in the context of the development of the Internet, where patient privacy information is threatened, and the prior art is difficult to achieve effective data sharing and model training while protecting privacy.
Through federated learning, a global detection model is jointly trained and deployed locally on each hospital client. The matching hash value is used to align data samples, crop and add noise to process model parameters to realize abnormal diagnostic behavior detection of patient diagnosis and treatment data.
It realizes abnormal behavior detection of patient diagnosis and treatment data, protects patient privacy, improves the efficiency and fairness of hospital diagnosis and treatment, and enhances the security of data privacy protection.
Smart Images

Figure CN120221104A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data security technology, and in particular, to a method, device, server, and computer storage medium for detecting abnormal diagnosis and treatment behaviors. Background Art
[0002] Currently, there are abnormal diagnosis behaviors in the domestic medical industry. For example, the reselling of expert numbers in large tertiary hospitals. Such abnormal diagnosis behaviors not only disrupt the hospital diagnosis and treatment system, but also undermine the concept of diagnosis and treatment fairness and reduce the efficiency of hospital diagnosis and treatment. Before the development of the Internet, the reselling of expert numbers was usually carried out by abnormal patients who purchased expert numbers by queuing in advance for abnormal diagnosis. With the development and application of Internet technology, abnormal patients generally collect the identity information and medical needs of other patients in advance on social platforms, bind the patient identity information first to violently grab numbers, and then unbind the identity information and mobile phone numbers after grabbing the numbers; or use other identity information to grab numbers and then agree with the patients who really need to register for a refund transaction at the time with the least transaction volume.
[0003] With the explosive growth of the number of users on the Internet, the threat to users' privacy information is increasing. The General Data Protection Regulation (GDPR) in the European Union and the Health Insurance Portability and Accountability Act (HIPAA) in the United States have clearly defined strict patient medical data protection regulations. In order to protect patients' privacy, Federated Learning, based on distributed machine learning, ensures that each hospital does not directly exchange or obtain the content of its respective database when jointly training a model, protecting patients' diagnosis and treatment information to a certain extent. Federated Learning is a cross-domain and distributed privacy protection method that allows multiple client databases to jointly train a model. It enables each client to train its own model locally and then send the model parameters to a centralized server for aggregation and then distribute them to each client, ensuring that all clients and the server can train models locally without exchanging the user information in the database, thus ensuring that attackers cannot directly obtain users' sensitive information. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides a method, device, server, and computer storage medium for detecting abnormal diagnosis and treatment behaviors, which jointly train a global detection model based on Federated Learning and can be deployed locally on each hospital client to achieve the detection of abnormal diagnosis behaviors in patients' diagnosis and treatment data.
[0005] The first aspect of the present application provides an abnormal diagnosis and treatment behavior detection method, and the method includes: Obtain a matching hash value and send the matching hash value to each hospital client among multiple hospital clients, so that each hospital client determines an aligned data sample based on the matching hash value, and the matching hash value is obtained from the hash value of each hospital client; Obtain the target model parameters of each hospital client to aggregate the target model parameters to obtain aggregated parameters. The target model parameters are obtained by performing a pruning process and a noise addition process on the original model parameters, and the original model parameters are obtained by each hospital client training a preset global model based on the aligned data sample; Send the aggregated parameters to each hospital client, so that each hospital client updates the target model parameters based on the aggregated parameters until a preset iteration termination condition is satisfied, and determine the global model that satisfies the iteration termination condition as the local detection model; Send the local detection model to each hospital client, so that when each hospital client receives patient diagnosis and treatment data, input the patient diagnosis and treatment data into the local detection model to determine whether there is an abnormal diagnosis behavior in the patient diagnosis and treatment data through the local detection model.
[0006] In an alternative embodiment, the pruning process on the original model parameters includes: Perform a pruning process on the original model parameters through the following pruning formula: ; where is the time step t is the i th original model parameter, C is a preset pruning threshold.
[0007] In an alternative embodiment, the noise addition process on the original model parameters includes: Calculate the maximum sensitivity of adjacent data sets, where the adjacent data sets are two data sets that differ by one diagnosis and treatment data record; Calculate the noise scale required for the uplink channel according to the privacy protection degree parameter and the maximum sensitivity, and the privacy protection degree parameter is determined by the selected differential privacy framework; Add the noise scale to the original model parameters after the pruning process in the uplink channel.
[0008] In an alternative embodiment, the calculating the noise scale required for the uplink channel according to the privacy protection degree parameter and the maximum sensitivity includes: Determine the noise scale through the following formula: ; where is the noise scale, P is the number of times of attack leakage per average aggregation operation in the uplink channel, is the maximum sensitivity, c is the control relaxation factor parameter in the differential privacy Gaussian mechanism parameter, is the privacy protection degree parameter.
[0009] In an optional implementation manner, calculating the maximum sensitivity of adjacent data sets includes: Determine the maximum sensitivity through the following formula: ; where is the sensitivity of the data set , m is the minimum value among the database sizes corresponding to the multiple hospital clients; Determine the sensitivity of the data set through the following formula: ; where is the parameter trained on the database , is the parameter trained on the adjacent database that only differs from the database by one piece of data , and are adjacent data sets.
[0010] In an optional implementation manner, aggregating the target model parameters to obtain aggregation parameters includes: Determine the number of records in the local database corresponding to each hospital client, and the total number of records of the multiple hospital clients; Determine the aggregation weight corresponding to each client according to the number of records and the total number of records; Perform weighted summation according to the aggregation weight and the target model parameters to obtain the aggregation parameters.
[0011] In an optional implementation manner, each hospital client updates the target model parameters based on the aggregation parameters, including: Determine the regularization term, and add the regularization term to the loss function of the global model to obtain the target loss function; Perform forward propagation on the aggregation parameters and the local database of the target hospital client to calculate the model output of the global model, where the target hospital client is any one of the multiple hospital clients; Calculate the loss value of the target loss function according to the model output and the true label of the local database; Calculate the model gradient according to the loss value, and update the target model parameters according to the model gradient.
[0012] A second aspect of the present application provides an abnormal diagnosis and treatment behavior detection device, the device includes: An acquisition module, configured to acquire a matching hash value, and send the matching hash value to each hospital client among the multiple hospital clients, so that each hospital client determines an aligned data sample based on the matching hash value, and the matching hash value is obtained through the hash value of each hospital client; An aggregation module, configured to acquire the target model parameters of each hospital client, and aggregate the target model parameters to obtain aggregation parameters, where the target model parameters are obtained by performing clipping processing and noise addition processing on the original model parameters, and the original model parameters are obtained by each hospital client training a preset global model based on the aligned data samples; A sending module, configured to send the aggregation parameters to each hospital client, so that each hospital client updates the target model parameters based on the aggregation parameters until a preset iteration termination condition is met, and determine the global model that meets the iteration termination condition as the local detection model; A detection module, configured to send the local detection model to each hospital client, so that when each hospital client receives patient diagnosis and treatment data, input the patient diagnosis and treatment data into the local detection model to determine whether there is an abnormal diagnosis behavior in the patient diagnosis and treatment data through the local detection model.
[0013] A third aspect of the present application provides a server, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps of the abnormal diagnosis and treatment behavior detection method are implemented.
[0014] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above abnormal diagnosis and treatment behavior detection method are implemented.
[0015] The abnormal diagnosis and treatment behavior detection method, device, server, and computer storage medium provided by the present application have at least one of the following beneficial effects: 1. By generating matching hash values, it can ensure that each hospital client can determine the aligned data samples based on the same benchmark (i.e., the hash value), thus solving the problem of inability to directly compare or jointly analyze due to data inconsistency; 2. Through the method of federated learning, each hospital client trains the model locally and only exchanges model parameters instead of the original data, thus protecting patient privacy and achieving model training while protecting the privacy of patient medical data; 3. The target model parameters are obtained by performing pruning processing and noise addition processing on the original model parameters, further enhancing privacy protection. Among them, pruning processing can reduce redundant information in the model parameters, and noise addition increases the difficulty for attackers to recover the original data from the model parameters; 4. In the framework of federated learning, the iterative process allows the model to learn from the data of multiple hospitals, thus capturing broader and more accurate features. When new patient diagnosis and treatment data is input into the model, the patient diagnosis and treatment data is detected in real time to determine whether there are abnormal diagnosis behaviors, improving the efficiency and fairness of hospital diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic structural diagram of an abnormal diagnosis and treatment behavior detection system shown in an embodiment of the present application; Figure 2 is another schematic structural diagram of an abnormal diagnosis and treatment behavior detection system shown in an embodiment of the present application; Figure 3 is a schematic flowchart of an abnormal diagnosis and treatment behavior detection method shown in an embodiment of the present application; Figure 4 is a schematic flowchart of a vertical federated detection method based on differential privacy shown in an embodiment of the present application; Figure 5 is a schematic diagram of a federated learning process shown in an embodiment of the present application; Figure 6 is a schematic diagram of a differential privacy mechanism shown in an embodiment of the present application; Figure 7 is a functional module diagram of an abnormal diagnosis and treatment behavior detection device shown in an embodiment of the present application; Figure 8 is a schematic structural diagram of a server shown in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0018] The concept, specific structure, and technical effects of the present invention will be clearly and completely described below in conjunction with the embodiments and the drawings to fully understand the purpose, features, and effects of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative efforts shall fall within the scope of protection of the present invention. In addition, all connection / linkage relationships involved in the patent do not simply refer to the direct connection of components, but rather refer to the formation of a more optimal connection structure by adding or reducing connection accessories according to specific implementation situations. The various technical features in the present invention can be interactively combined without conflicting with each other.
[0019] Referring to Figure 1 As shown, the abnormal diagnosis and treatment behavior detection system may include multiple hospital clients. The multiple hospital clients include a first hospital client and a second hospital client. The first hospital client is any one of the multiple hospital clients and serves as a server for terminal processing. The second hospital client is all the hospital clients other than the first hospital client among the multiple hospital clients. By initializing the first hospital client and broadcasting and sending the global model to the second hospital client, then both the first hospital client and the second hospital client train the global model based on their local databases. After the second hospital client finishes training, it sends the trained second model parameters to the first hospital client so that the first hospital client performs aggregation processing based on the first model parameters obtained by its own training and the second model parameters to obtain aggregation parameters. Further, the first hospital client sends the aggregation parameters to the second hospital client so that the first hospital client updates the first model parameters based on the aggregation parameters, and the second hospital client updates the second model parameters based on the aggregation parameters to obtain a local detection model corresponding to each hospital client. When receiving newly input patient diagnosis and treatment data, the first hospital client and / or the second hospital client can judge the patient diagnosis data based on their own local detection models to determine whether there are abnormal diagnosis and treatment behaviors in the patient diagnosis data. Referring to Figure 2 As shown, in addition to including multiple hospital clients, the abnormal diagnosis and treatment behavior detection system may further include a server. The server replaces the initialization of the first hospital client and broadcasts and sends the global model to each hospital client among the multiple hospital clients. Among them, the server can be designated by multiple hospital clients. When each hospital client receives the global model sent by the server, each hospital client can train the global model based on its own corresponding local database to obtain model parameters, and send the model parameters to the server so that the server aggregates the model parameters to obtain aggregation parameters.
[0020] Referring toFigure 3 As shown in the figure, it is a schematic flowchart of an abnormal diagnosis and treatment behavior detection method shown in an embodiment of the present application. The abnormal diagnosis and treatment behavior detection method includes the following steps.
[0021] S31. Obtain a matching hash value and send the matching hash value to each hospital client among a plurality of hospital clients, so that each hospital client determines an aligned data sample based on the matching hash value.
[0022] Wherein, the matching hash value is obtained from the hash values of each hospital client.
[0023] Since vertical federated learning requires a large overlap in the training samples of participants and a small overlap in data features, data sample alignment can be used to find the common samples owned by participants (i.e., the overlapping part of the training samples) for subsequent joint learning of feature dimensions. In some embodiments, each hospital client can first encrypt the local patient data in the local database using an asymmetric encryption algorithm (e.g., Bind RSA) so that even if the subsequent model parameters are intercepted during transmission, they cannot be decrypted and understood by unauthorized third parties. Then, the encrypted local patient data is converted into a hash value using a hash algorithm. The selection of the hash function should ensure that even very similar data will produce completely different hash values, thus protecting data privacy. Specifically, the SHA-256 algorithm can be used to map the original data to a string similar to a 256-bit data "fingerprint", and this mapping is irreversible. Through operations in the main part of the SHA-256 algorithm, such as multiple non-linear functions and bit operations, where non-linearity means that a small change in the input will not linearly affect the output but will cause a huge difference in the output, it is ensured that even a small difference in the input data will be amplified during the iteration process. Due to the one-way nature of the hash function, sharing the hash value itself does not disclose any data information of the local patient data. Each hospital client shares its respective hash value with the server so that the server can determine the matching hash values among multiple hash values. That is, after receiving the hash values from each hospital client, the server queries for the same hash values among multiple hash values and then determines the hash values from different hospital clients that are the same as the matching hash values. When the matching hash values are determined, the server distributes the matching hash values to each hospital client so that each hospital client can determine which data samples are aligned based on the matching hash values. The local patient data corresponding to the matching hash values is the aligned data sample, which is called the aligned data sample. The local patient data may include high-identification-rate information such as the patient's ID number and phone number, as well as diagnosis and treatment information such as the registration department and the number of times of registering with the same doctor in the past that are required for discrimination. It should be noted that the selection of the local patient data used to generate the hash value is based on its importance and sensitivity during the alignment process to ensure both effective data alignment and maximum protection of patient privacy.
[0024] Through the above optional implementation methods, the secure alignment of data samples between hospitals is achieved through the hash algorithm, protecting patient privacy and ensuring the effective alignment of data samples. Based on the aligned data samples, each hospital can conduct joint learning, jointly train the global model, improve the model performance, and be used for the detection of abnormal diagnosis and treatment behaviors.
[0025] S32, Obtain the target model parameters of each hospital client to aggregate the target model parameters to obtain aggregated parameters.
[0026] Among them, the target model parameters are obtained by performing pruning processing and noise addition processing on the original model parameters, and the original model parameters are obtained by each hospital client training a preset global model based on the aligned data samples.
[0027] When each hospital client obtains the global model and the corresponding model parameters sent by the server, each hospital client can perform forward propagation based on the original model parameters and its corresponding local database, calculate the model output of the global model, calculate the loss value of the loss function of the global model through the root model output and the true labels in the local database, then calculate the model gradient according to the loss value, and use the modified model gradient as the updated model parameters. It should be noted that in the embodiments of the present application, the model parameters obtained after the first iterative training are referred to as the original model parameters. When multiple iterative trainings are performed subsequently, the model parameters obtained from the previous round of training are referred to as the original model parameters. After obtaining the original model parameters, each hospital client needs to upload its corresponding original model parameters to the server, so that the server aggregates the original model parameters of each hospital client to obtain aggregation parameters, and distributes the aggregation parameters to each hospital client, so that each hospital client enters the next round of iterative training. That is, after the data samples are aligned, each hospital client can perform joint learning based on the aligned data samples to improve the performance of the global model, and by sharing the model parameters (where the model parameters include gradient information), each hospital client can jointly train a global model for the detection of abnormal diagnosis and treatment behaviors.
[0028] According to the differences in the data space and feature space, federated learning can include horizontal federated learning, vertical federated learning, and federated transfer learning. For abnormal patient detection, different hospitals in the same region are good at different diagnosis and treatment directions, so there will not be too much overlap in the data feature spaces of different hospitals. Therefore, vertical federated learning is more suitable for abnormal patient detection. In order to combine the data with different feature spaces and the same sample space of different hospitals, vertical federated learning is used to learn the data features. In vertical federated learning, each party has its own local database and local global model, and cross-domain learning is achieved by exchanging intermediate results. Referring together Figure 4 In the embodiments of the present application, the server needs to initialize a global model and send the global model and the corresponding model parameters to each hospital client. Specifically, the server can first construct a global model based on a multi-layer perception network with 256 hidden layers, where the global model uses ReLU units and Sigmoid function (used to generate the binary classification probability of whether it is an abnormal diagnosis and treatment behavior) units. Specifically, the loss function is the cross-entropy function, expressed as wherey denotes the true label on the test set, p represents the probability that the sample predicted by the global model is an abnormal diagnosis and treatment behavior. To avoid overfitting during model training, a penalty term related to the size of the model parameters is added to the loss function as a regularization term to limit the value range of the model parameters, thereby preventing the model from being too complex and reducing the risk of overfitting. Among them, the regular expression corresponding to the regularization term is , is the model parameter of the previous round, is the model parameter of the current round of the th hospital client, represents a hyperparameter that can be adjusted by different hospital clients according to their actual situations, and is used to prevent the model from overfitting the training parameters of a certain round. Specifically, during model training, an expression containing is added after the loss function, expressed as , and the regularization term is used to encourage the model parameter of the current round not to deviate too much from the global model parameter of the previous round. Too much.
[0029] In an optional embodiment, the noise addition process for the original model parameters includes: Calculating the maximum sensitivity of adjacent data sets, where the adjacent data sets are two data sets that differ by one diagnosis and treatment data record; Calculating the noise scale required for the uplink channel according to the privacy protection degree parameter and the maximum sensitivity, where the privacy protection degree parameter is determined by the selected differential privacy framework; Adding the noise scale to the original model parameters after the clipping process in the uplink channel.
[0030] Since the model parameters transmitted between each hospital client and the server in federated learning may disclose privacy information, in the embodiments of the present application, differential privacy technology is used to desensitize the transmitted model parameters. Among them, differential privacy technology provides a standardized measure for the privacy protection level of data analysis, and it quantifies the impact of the change of a single data point on the algorithm result from a statistical perspective. Referring together to Figure 5 and Figure 6 , in the embodiments of the present application, The differential privacy framework sets a clear privacy protection criterion for the distributed data system, where and are both privacy protection degree parameters, is a parameter greater than 0, used to measure the privacy protection level, The smaller the value, the stronger the privacy protection. The parameters describe that when adopting level of protection measures, two adjacent data sets and at probability may not be protected by differential privacy. Further, by defining a distance function to calculate the maximum sensitivity on adjacent data sets, where the maximum sensitivity reflects the maximum impact of the change of a single data point in the data set on the algorithm result, and adjacent data sets are two data sets that differ only by one medical record, the purpose is that the attacker cannot discover the information of the different record. By defining differential privacy as , where Pr is the probability value, indicating that the adjacent data sets and are transformed into the same output range by the mechanism, M is the introduced Gaussian mechanism algorithm, and S represents the set of all possible output results. To make the added Gaussian noise distribution meet the requirements of the differential privacy mechanism, it is necessary to set the standard deviation of the random noise distribution to meet , where the parameter meets , is the maximum sensitivity of the adjacent database, defined as , represents the smallest unit of private data to be protected in differential privacy, and represents the model parameter in the embodiment of this application, that is = represents the parameter obtained by training on the database , represents the parameter obtained by training on an adjacent database that differs from the database by only one piece of data , s are the model parameters of each hospital client.
[0031] In the embodiment of this application, for the noise scale to be added to the model parameters of each hospital client, when determining the degree of privacy protection of two parameters, the parameters to be determined also include the maximum sensitivity on the noise to be added to the uplink channel, that is, the maximum sensitivity of the adjacent data sets. Specifically, the server can determine the maximum sensitivity of the adjacent data sets through the following formula : ; Among them, is the sensitivity of the data set , mis the minimum value among the database sizes corresponding to the multiple hospital clients.
[0032] Determine the dataset through the following formula sensitivity : ; where is , is , obtained by each hospital client based on training with the local database, where , and are adjacent datasets.
[0033] After determining the maximum sensitivity of adjacent databases, each hospital client can determine the noise scale to be added on the uplink channel according to the maximum sensitivity and the privacy protection degree parameter of the differential privacy technology. Specifically, assume that in the uplink channel, on average, each aggregation operation will be leaked P times due to attacks. Due to the linear relationship between the noise scale and the privacy budget , the noise added on the uplink channel should be Gaussian distribution noise with a scale of where c is the control relaxation factor parameter in the differential privacy Gaussian mechanism , and satisfies .
[0034] Since the model parameters generally do not have a fixed upper bound, it is necessary to agree on the upper bound C of the parameters. Therefore, before adding noise, it is necessary to clip the parameters to ensure that the parameter values are within a controllable range. The specific clipping method is , and ensure , where C is the preset clipping threshold, so as to obtain the sensitivity on the dataset is , and further the maximum sensitivity of the differential privacy mechanism added on the uplink channel can be obtained as , where is the minimum value among the database sizes of all clients.
[0035] Through the above optional implementation manners, by introducing differential privacy technology, the problem that privacy information may be leaked during the transmission of model parameters between the hospital client and the server in federated learning is effectively solved. First, the maximum sensitivity of adjacent data sets is calculated, and then the noise scale required for the uplink channel is determined according to the privacy protection level parameter and the maximum sensitivity, and this noise is added to the original model parameters after clipping processing, which not only provides a clear privacy protection criterion for the distributed data system, but also ensures the privacy and security of the model parameters through reasonable noise addition and clipping processing, while maintaining the usability of the model.
[0036] S33, sending the aggregated parameters to each hospital client, so that each hospital client updates the target model parameters based on the aggregated parameters until a preset iteration termination condition is met, and determining the global model that meets the iteration termination condition as the local detection model.
[0037] After the global model is constructed, the server sends the global model and the corresponding initial training parameters to each hospital client, so that each hospital client trains the global model according to the local database to obtain the corresponding model parameters. The initial training parameters include initial model parameters , preset parameters in the regular expression , and the preset number of iteration rounds required to ensure that the global model converges to a certain extent , and in the differential privacy protection level parameters and parameters. When each hospital client receives the global model, it trains the global model based on the local database until a preset iteration termination condition is met, and outputs the trained model parameters. For example, K hospital clients each include their corresponding local data sets, and all data sets are represented as , each hospital client initializes the local model parameters , and in each communication round, each hospital client i processes its corresponding batch data . Then, each hospital client can upload its corresponding model parameters to the server, so that the server can aggregate the model parameters of each hospital client to obtain aggregated parameters.
[0038] In an optional implementation manner, the aggregating the target model parameters to obtain aggregated parameters includes: Determining the number of records in the local database corresponding to each hospital client and the total number of records of the multiple hospital clients; Determining the aggregation weight corresponding to each client according to the number of records and the total number of records; Perform a weighted sum according to the aggregation weights and the target model parameters to obtain the aggregation parameters.
[0039] In some embodiments, the server may, according to aggregate the model parameters of all hospital clients, where is the aggregation weight, and N is the number of hospital clients. Among them, the aggregation weight , is the number of records in the local database of the th hospital client, and is the total number of records of all hospital clients. Further, the server may send the aggregation parameters to each hospital client through the downlink channel, so that each hospital client updates the local model parameters based on the aggregation parameters, that is, further trains the local global model based on the aggregation parameters. During the update process, each participating client verifies whether the performance of the global model meets the requirements after each iteration training. If the preset accuracy requirement is reached, the iteration process ends or a single hospital client temporarily exits the iteration process to obtain a local detection model; if not, the next round of local training is performed until the preset accuracy requirement is met. Specifically, each hospital client updates the target model parameters based on the above-mentioned aggregation parameters, including the following steps: (1) Initialize the model parameters.
[0040] Before the first round of training, the hospital client usually receives the global model parameters, that is, the aggregation parameters, from the server as the initial parameters. For each round of local training, the hospital client will use the global model parameters of the previous round as the starting point.
[0041] (2) Define the loss function.
[0042] According to the same embodiment as in step S32, the loss function is usually calculated based on the local dataset of the hospital client i , for example, the cross-entropy loss, expressed as , and a regularization term is added.
[0043] (3) Forward propagation.
[0044] During the training process, use the current model parameters and the local dataset of the hospital client i to perform forward propagation and calculate the model output.
[0045] (4) Calculate the loss.
[0046] Calculate the loss function according to the model output and the true label, and calculate the regularization term, and add the two to obtain the total loss value.
[0047] (5) Backpropagation.
[0048] Calculate the model gradient based on the total loss value, i.e., the derivative of the loss function with respect to the model parameters, including the gradient of the loss function with respect to the current model parameters and the gradient of the regularization term with respect to the current model parameters.
[0049] (6) Parameter update.
[0050] Use an optimizer (such as SGD, Adam, etc.) to update the model parameters according to the gradient, where the optimizer can adjust the update step size of the parameters according to the learning rate.
[0051] (7) Iterative training.
[0052] Repeat steps 3 to 6 until the stopping condition is met (such as reaching the maximum number of iterations, loss convergence, etc.). Among them, when the hospital client trains the model locally, it converges the global model according to the method.
[0053] (8) Upload model parameters.
[0054] After the training is completed, each hospital client uploads the model parameters obtained from local training to the server, so that the server aggregates the model parameters received from each hospital client to update the global model parameters, and distributes the updated global model parameters to each hospital client as the starting point for the next round of local training.
[0055] Through the above optional implementation manners, by determining the aggregation weight according to the number of records in the local database of the hospital client, the weighted aggregation of the target model parameters is realized, improving the accuracy and generalization ability of the global model. At the same time, through iterative training and parameter update, the local model is continuously optimized, accelerating the convergence of the global model and enhancing the efficiency and effect of federated learning in medical data sharing.
[0056] In other embodiments, a specific guest client (i.e., the first hospital client) can be preset in advance. Each hospital client calculates the output of the local global model , and sends all of them to the guest client, so that the guest client aggregates the indices of the local global models sent by all hospital clients z to obtain , where is the activation function, and calculates the prediction output locally and the loss function between the true label . Further, the guest client calculates the loss function with respect to . The gradient value, and send this gradient value as an aggregation parameter to each hospital client, so that each hospital client i Updates the local model parameters according to the received aggregation parameter, specifically: ; Among them, Is the learning rate in model training.
[0057] Also refer to Figure 5 And Figure 6 , in an alternative embodiment, since the model parameters transmitted between each hospital client and the server may disclose privacy information, before the server sends the aggregation parameter to each hospital client, in order to ensure that the downlink channel for the server to send the aggregation parameter is also protected by the corresponding differential privacy mechanism, according to the same embodiment of step S32, it is also necessary to perform noise-added desensitization protection on the parameters sent by the server starting from defining the maximum sensitivity. In the downlink channel, define the distance function . Similarly, the maximum sensitivity on adjacent data sets is . Since all other data sets in the local data set of the hospital client are exactly the same, only the adjacent data sets And Are different, so only the difference between Needs to be compared.
[0058] In the embodiment of the present application, the uplink channel and the downlink channel are combined as an integrated aggregation process, so the differential privacy mechanisms of the uplink channel and the downlink channel need to be combined into a global Gaussian mechanism. Therefore, it is necessary to consider the distribution of all the noises added to the uplink channel. The aggregated parameter obtained by the server after adding noises by each hospital client is: ; Since the noises added on each hospital client are independently and identically distributed, then for the overall noise Of the distribution Can get , then the noise scale of the distribution Is determined as , Is the number of hospital clients. For the global model, it is set that initially Rounds are required to obtain the best convergence performance of the model. Then the noise scale of the global differential privacy mechanism should be . Also, since the variance values of the noises in the uplink channel and the downlink channel are added to obtain the variance value of the global noise, then the noise value of the downlink channel should be: ; That is, before the server sends the global parameter to each hospital, it is necessary to add a noise scale of of Gaussian distributed noise to ensure that the privacy protection level is protected by the differential privacy mechanism of
[0059] When determining the uplink channel noise scale and the downlink channel noise scale After that, the global noise scale is determined by the following formula : ; Further prove whether the model converges after adding Gaussian noise with a scale of to the model parameters. Specifically, the conclusion can be obtained through the method of mathematical calculation and proof. Generally, the convergence speed of the model is a convex function of the aggregation times T, that is, there exists an optimal T that can make the model converge fastest. After obtaining the parameter T, the value of T can be adjusted, or the experimental method can be used to approximate the optimal range of T.
[0060] It should be noted that the differential privacy algorithm of the Gaussian mechanism is used once for the uplink channel and the downlink channel respectively. In the uplink channel, each hospital client only needs to protect its own data information, while the server needs to protect all aggregated information. Although the same sensitivity calculation method is used, the calculated sensitivities are different.
[0061] Through the above optional implementation manners, by adding noise for desensitization protection to the aggregated parameters sent by the server in the downlink channel, it is ensured that the entire aggregation process (including the uplink channel and the downlink channel) is protected by the differential privacy mechanism. By defining the distance function and calculating the maximum sensitivity on adjacent data sets, combined with the noise distribution situation added to the uplink channel, the noise scale that needs to be added to the downlink channel is determined, effectively preventing the privacy leakage of model parameters during the transmission process, improving the security and privacy protection level of federated learning, and at the same time ensuring the convergence performance and accuracy of the global model.
[0062] S34. Send the local detection model to each hospital client so that when each hospital client receives patient diagnosis and treatment data, input the patient diagnosis and treatment data into the local detection model to determine whether there is an abnormal diagnosis behavior in the patient diagnosis and treatment data through the local detection model.
[0063] After iterating to a certain number of communication rounds, each hospital client has a local model locally that can be used to determine whether new medical treatment data is abnormal. And if the data samples of each hospital change, the model can be made to fit the existing data by repeating the iteration, and all iteration processes are protected by the differential privacy algorithm, ensuring that the attacker cannot obtain valid information by sending the differences between the data on the uplink and downlink channels. In some embodiments, when the target hospital client receives new patient medical treatment data, the patient medical treatment data is input into its corresponding local detection model, and the target hospital client is any one of the multiple hospital clients. The local detection model processes the patient medical treatment data according to the updated model parameters and determines whether there is any abnormal medical treatment behavior in the new patient medical treatment data based on the processing result. When the data features do not match the features learned in the local detection model, or the data exceeds the normal range of the local detection model, the local detection model can output 0, indicating that there is abnormal medical treatment behavior in the patient medical treatment data, otherwise the local detection model can output 1, indicating that there is no abnormal medical treatment behavior in the patient medical treatment data. When it is determined that there is abnormal medical treatment behavior, the target hospital client can trigger an alarm prompt.
[0064] Refer to Figure 7 As shown, it is a functional module diagram of the abnormal medical treatment behavior detection device shown in the embodiments of the present application.
[0065] In some embodiments, the abnormal medical treatment behavior detection device 70 may include multiple functional modules composed of computer program segments. The computer programs of each program segment of the abnormal medical treatment behavior detection device 70 can be stored in the memory of the server and executed by at least one processor to perform the functions of detecting abnormal medical treatment behavior (see details in Figure 3 description). According to the functions it performs, it can be divided into multiple functional modules. The functional modules may include: an acquisition module 701, an aggregation module 702, a sending module 703, and a detection module 704. The module referred to in the present application means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0066] The acquisition module 701 is used to acquire a matching hash value and send the matching hash value to each hospital client among the multiple hospital clients, so that each hospital client determines the aligned data samples based on the matching hash value, and the matching hash value is obtained through the hash value of each hospital client.
[0067] The aggregation module 702 is configured to obtain the target model parameters of each hospital client, aggregate the target model parameters to obtain aggregation parameters. The target model parameters are obtained by performing a clipping process and a noise addition process on the original model parameters, and the original model parameters are obtained by each hospital client training a preset global model based on the aligned data samples.
[0068] The sending module 703 is configured to send the aggregation parameters to each hospital client, so that each hospital client updates the target model parameters based on the aggregation parameters until a preset iteration termination condition is met, and determines the global model that meets the iteration termination condition as the local detection model.
[0069] The detection module 704 is configured to send the local detection model to each hospital client, so that when each hospital client receives patient diagnosis and treatment data, it inputs the patient diagnosis and treatment data into the local detection model to determine whether there is an abnormal diagnosis behavior in the patient diagnosis and treatment data through the local detection model.
[0070] It should be understood that the various change modes and specific embodiments in the abnormal diagnosis and treatment behavior detection method provided in the above embodiments are equally applicable to the abnormal diagnosis and treatment behavior detection device in this embodiment. Through the foregoing detailed description of the abnormal diagnosis and treatment behavior detection method, those skilled in the art can clearly know the implementation method of the abnormal diagnosis and treatment behavior detection device in this embodiment. For the sake of brevity of the specification, it will not be described in detail here.
[0071] Refer to Figure 8 As shown in the structural schematic diagram of the server shown in the embodiment of the present application. In a preferred embodiment of the present application, the server 8 includes a memory 81, at least one processor 82, and at least one communication bus 83.
[0072] Those skilled in the art should understand that Figure 8 the structure of the server shown does not constitute a limitation on the embodiment of the present application. It can be a bus structure or a star structure. The server 8 may further include more or fewer other hardware or software than shown, or different component arrangements.
[0073] In some embodiments, the server 8 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application specific integrated circuits, programmable gate arrays, digital signal processors, and embedded devices, etc. The server 8 may further include a hospital client device, and the hospital client device includes, but is not limited to, any electronic product that can perform human-computer interaction with a user through means such as a keyboard, mouse, remote control, touchpad, or voice control device. For example, a computer, a tablet computer, a smart phone, a digital camera, etc.
[0074] In the above embodiments provided by the present application, it should be understood that the disclosed methods, devices, computer-readable storage media, and servers can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple components or modules can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or components or modules can be in electrical, mechanical, or other forms.
[0075] The components described as separate components may or may not be physically separated, and the components displayed as components may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the components can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0076] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each component can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0077] When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0078] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily all essential to the present invention.
[0079] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0080] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A method for detecting abnormal diagnosis and treatment behavior, characterized in that: The method comprises: Obtaining a matching hash value, and sending the matching hash value to each of the multiple hospital clients, so that each of the hospital clients determines an alignment data sample based on the matching hash value, wherein the matching hash value is obtained by the hash value of each of the hospital clients; Obtaining target model parameters of each hospital client to aggregate the target model parameters to obtain aggregated parameters, wherein the target model parameters are obtained by performing a trimming process and a noise adding process on the original model parameters, and the original model parameters are obtained by each hospital client training a preset global model based on the aligned data samples; Sending the aggregation parameters to each of the hospital clients, so that each of the hospital clients updates the target model parameters based on the aggregation parameters until a preset iteration termination condition is met, and determining the global model that meets the iteration termination condition as the local detection model; The local detection model is sent to each hospital client so that when each hospital client receives the patient's diagnosis and treatment data, the patient's diagnosis and treatment data is input into the local detection model so as to determine whether the patient's diagnosis and treatment data has abnormal diagnostic behavior through the local detection model.
2. The abnormal diagnosis and treatment behavior detection method according to claim 1, characterized in that: The clipping process of the original model parameters includes: The original model parameters are clipped using the following clipping formula: ; in, is the time step t No. i The original model parameters, C is the preset clipping threshold.
3. The abnormal diagnosis and treatment behavior detection method according to claim 2, characterized in that: The noise adding process to the original model parameters comprises: Calculating the maximum sensitivity of adjacent data sets, where the adjacent data sets are two data sets that differ by one diagnosis and treatment data record; Calculating a noise scale required for an uplink channel according to a privacy protection degree parameter and the maximum sensitivity, wherein the privacy protection degree parameter is determined by a selected differential privacy framework; The noise scale is added to the original model parameters after clipping in the uplink channel.
4. The abnormal diagnosis and treatment behavior detection method according to claim 3, characterized in that: The noise scale required for calculating the uplink channel according to the privacy protection degree parameter and the maximum sensitivity includes: The noise scale is determined by the following formula: ; in, is the noise scale, P is the average number of attack leaks in each aggregation operation in the uplink channel, is the maximum sensitivity, c is the control relaxation factor in the differential privacy Gaussian mechanism Parameters, is the privacy protection degree parameter.
5. The abnormal diagnosis and treatment behavior detection method according to claim 3, characterized in that: The calculating the maximum sensitivity of the adjacent data sets includes: The maximum sensitivity is determined by the following formula: ; in, For the dataset sensitivity, m is the minimum value among the database sizes corresponding to the multiple hospital clients; The data set is determined by the following formula: Sensitivity: ; in, For the database The parameters obtained by training above, For the database On adjacent databases that differ by only one piece of data The parameters obtained by training, and is the adjacent data set.
6. The abnormal diagnosis and treatment behavior detection method according to claim 1, characterized in that: The aggregating the target model parameters to obtain aggregated parameters comprises: Determine the number of records in the local database corresponding to each hospital client and the total number of records of the multiple hospital clients; Determine the aggregation weight corresponding to each client according to the number of records and the total number of records; The aggregation parameter is obtained by performing weighted summation on the aggregation weight and the target model parameter.
7. The abnormal diagnosis and treatment behavior detection method according to claim 1, characterized in that: The updating of the target model parameters by each hospital client based on the aggregation parameters includes: Determining a regularization term, and adding the regularization term to the loss function of the global model to obtain a target loss function; Forward propagation of the aggregation parameters and the local database of the target hospital client to calculate the model output of the global model, where the target hospital client is any one of the multiple hospital clients; Calculate the loss value of the target loss function according to the model output and the true label of the local database; A model gradient is calculated according to the loss value, and the target model parameters are updated according to the model gradient.
8. An abnormal diagnosis and treatment behavior detection device, characterized in that: The device comprises: An acquisition module, configured to acquire a matching hash value, and send the matching hash value to each of a plurality of hospital clients, so that each of the hospital clients determines an alignment data sample based on the matching hash value, wherein the matching hash value is acquired by a hash value of each of the hospital clients; an aggregation module, used for acquiring target model parameters of each hospital client, and aggregating the target model parameters to obtain aggregated parameters, wherein the target model parameters are obtained by performing a clipping process and a noise adding process on the original model parameters, and the original model parameters are obtained by each hospital client training a preset global model based on the aligned data samples; A sending module, used for sending the aggregation parameters to each of the hospital clients, so that each of the hospital clients updates the target model parameters based on the aggregation parameters until a preset iteration termination condition is met, and determines the global model that meets the iteration termination condition as a local detection model; The detection module is used to send the local detection model to each hospital client, so that when each hospital client receives the patient diagnosis and treatment data, the patient diagnosis and treatment data is input into the local detection model to determine whether the patient diagnosis and treatment data has abnormal diagnostic behavior through the local detection model.
9. A server, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the abnormal diagnosis and treatment behavior detection method described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the abnormal diagnosis and treatment behavior detection method described in any one of claims 1 to 7 are implemented.