Federal learning-based oral medical privacy data desensitization method and system

By adopting technologies such as federated learning and reinforcement learning in oral medical data processing, we dynamically adjust privacy parameters and introduce anomaly detection mechanism, we solve the privacy leakage problem during oral medical data processing and sharing, and achieve safe and efficient data collaboration and model performance improvement.

CN120030602AActive Publication Date: 2025-05-23THE 900TH HOSPITAL OF THE CHINESE PEOPLES LIBERATION ARMY JOINT LOGISTICS SUPPORT FORCE

Patent Information

Application Number
CN202510512004.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Due to its high privacy and diverse forms, oral medical data faces the risk of data leakage when processed and shared, and it is difficult to fully tap the value of data while protecting privacy.

Method used

The oral medical privacy data desensitization method based on federated learning is adopted, and a federated learning system is built by obtaining multi-source data for classification, cleaning and desensitization. The federated learning system is built, and the differential privacy parameters are dynamically adjusted using reinforcement learning algorithms, and an abnormal detection mechanism is introduced into the transmission gradient to improve the security of the system and model performance.

Benefits of technology

The collaborative optimization of privacy protection and diagnostic model performance is achieved, providing safe, efficient and sustainable solutions for cross-regional and multi-institutional medical collaboration, reducing the risk of privacy leakage caused by data flow, and improving the overall diagnostic capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030602A_ABST
    Figure CN120030602A_ABST
Patent Text Reader

Abstract

The invention relates to an oral medical privacy data desensitization method and system based on federal learning, and the method comprises the following steps: obtaining multi-source data related to oral medical treatment, and carrying out the classification, cleaning and desensitization of different types of medical data; based on the processed high-quality data, constructing a federal learning system to realize collaborative modeling across oral medical institutions; through intelligent dynamic adjustment of privacy parameters and attack defense, the security and model performance of the federal learning system are improved. According to the invention, collaborative optimization of privacy protection and diagnostic model performance can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data management, and in particular to a method and system for desensitizing oral medical privacy data based on federated learning. Background Art

[0002] With the rapid development of modern medical technology, especially the widespread application of artificial intelligence (AI) in the medical field, the oral medical industry is gradually moving towards digitalization and intelligence. However, as a kind of sensitive data with high privacy and specificity, oral medical data faces many technical and legal challenges in its processing and sharing: on the one hand, these data include oral images (such as dental X-rays, 3D modeling), clinical cases, treatment plans and other diverse forms, which are highly related to the patient's identity and are very easy to expose privacy; on the other hand, when sharing data between medical institutions or between scientific research institutions and medical institutions for model training or research, data leakage must be strictly prevented. Therefore, how to fully tap the value of these data while protecting data privacy has become an urgent problem to be solved. Summary of the invention

[0003] In order to solve the above problems, the purpose of the present invention is to provide an oral medical privacy data desensitization method and system based on federated learning, which can achieve the coordinated optimization of privacy protection and diagnostic model performance, and provide a safe, efficient and sustainable solution for cross-regional and multi-institutional medical collaboration.

[0004] To achieve the above object, the present invention adopts the following technical solutions: A method for desensitizing oral medical privacy data based on federated learning, comprising the following steps: Acquire multi-source data related to oral healthcare, and classify, clean and desensitize different types of medical data; Based on the processed high-quality data, a federated learning system is built to achieve collaborative modeling across dental medical institutions; By intelligently and dynamically adjusting privacy parameters and attack defense, the security and model performance of the federated learning system are improved as follows: Introducing a reinforcement learning algorithm to dynamically adjust differential privacy parameters and control the noise amplitude according to the current training round and model accuracy to minimize the impact on model performance; An anomaly detection mechanism is introduced in the transmission gradient to monitor whether the uploaded updates are abnormal and defend against data poisoning or gradient back-propagation attacks.

[0005] Furthermore, we obtain multi-source data related to oral healthcare, and classify, clean and desensitize different types of medical data, as follows: The medical data includes oral imaging data, structured data and unstructured text data; Check and standardize data formats, remove redundant fields and useless data; Oral image data uses a deep learning segmentation algorithm to separate the tooth area and facial information, removing unnecessary privacy features; Structured data processing involves hashing or replacing sensitive fields; Unstructured text data uses NLP models to automatically detect and mask sensitive information such as patient names and ID numbers.

[0006] Furthermore, the oral image data uses a deep learning segmentation algorithm to separate the tooth area and facial information and remove unnecessary privacy features, as follows: Obtain oral imaging data, including dental X-rays and CT slices, and standardize the oral imaging data to a fixed resolution to ensure that the deep learning model supports consistent size: ; in, The height and width corresponding to the target resolution; is the original image; is the resized image; is a pixel; is the size adjustment function; According to the professional annotation tool of stomatology, the segmentation mask of the tooth area in the image is manually annotated to form the training label M(x,y); Build a segmentation model based on the UNet model, which consists of an encoder and a decoder and supports high-precision pixel-level segmentation; Encoder: progressively downsamples the image and extracts multi-scale features using multiple layers of convolution and pooling; Decoder: progressively upsamples, uses skip connections to pass encoder features to the decoder, and combines high- and low-level feature information; According to the training data set, input the image and the true label mask M(x,y), calculate the loss L through forward propagation, and use the Adam optimizer to update the model parameters; Input a new unlabeled oral image, and output the tooth area segmentation result after the segmentation model predicts it. , Based on the tooth region segmentation results , the non-tooth area is set to black or pseudo data noise to generate a desensitized image.

[0007] Furthermore, a segmentation model is constructed based on the UNet model, as follows: In the encoder part, two convolution operations are performed consecutively: ; Among them, is the feature map, is the convolution kernel, is the bias, is a nonlinear activation function; For input data; Use max pooling to progressively downsample the feature map: ; Among them, P(X) represents the output feature map after downsampling; K=k*l is the pooling window size; is the maximum value of all values ​​in the pooling window; X is the input feature map; In the decoder part, the spatial resolution of the feature map is increased by upsampling: ; Among them, U(X) represents the output feature map after upsampling; Indicates upsampling; scale=2 is the upsampling multiple, which means that the resolution of the feature map in each dimension is increased by 2 times; Each layer of decoder is concatenated with the corresponding encoder feature map through skip connection: ; Among them, C is the concatenated feature map; Concat represents the concatenation operation; are the outputs of the decoder and encoder respectively; Output predicted segmentation mask , the predicted segmentation mask value is a probability distribution value between [0,1]; During training, a weighted combination of BCE and Dice is used as the loss function L: ; ; ; Among them, λ 1 ,λ 2 represents the weight of BCE and Dice Loss; L BCE is the loss of BCE; L Dice is the Dice loss, and N is the number of samples.

[0008] Furthermore, the federated learning system includes independent nodes and a central server, specifically: Treat each medical institution as an independent node, use the local cleaned data of the medical institution to participate in distributed training, upload only model parameters without sharing original data; the node local model processes specific tasks; Each node Using its local dataset , according to the current weight W tPerform local gradient updates :

[0009] in, η is the learning rate; For Node Loss function for medical tasks; represents the gradient function; Central server: schedules training tasks and collects model updates from all nodes After that, the global model is aggregated to generate the optimized global model weight W t+1 : ; Where n is the number of nodes; The optimized global model weights are distributed back to each node for the next round of iteration.

[0010] Furthermore, a reinforcement learning algorithm is introduced to dynamically adjust the differential privacy parameters and control the noise amplitude according to the current training round and model accuracy to minimize the impact on model performance. The details are as follows: a. Initialize the federated learning global model W 0 , Initial Privacy Budget , noise intensity σ 0 ; Start the reinforcement learning agent and generate privacy parameter adjustment actions according to the initial strategy; b. Each node Use the privacy budget at time t and the noise intensity σ t , in the local dataset The model training is completed on the GitHub repository; the implementation formula for differential privacy is: ; in, is the gradient update; Gaussian noise generated by the current noise intensity; For Node On-premises data The gradient of the loss function L on the model parameters; I is the identity matrix; represents differential privacy noise; c. The central server aggregates the updates uploaded by the nodes to obtain the global model; d. According to the current status , using reinforcement learning models to generate new privacy parameter adjustment actions ; Among them, L t is the loss function at the current moment; are the adjustment amounts of privacy budget and noise intensity respectively; Execute actions to update privacy parameters: ; Use RL algorithms to update reinforcement learning policy network weights; e. Privacy parameters Model weight W t+1 Distribute to nodes and enter the next round of training, and loop until the termination condition is met.

[0011] Furthermore, an anomaly detection mechanism is introduced in the transmission gradient to monitor whether the uploaded update has anomalies and defend against data poisoning or gradient back-pushing attacks, as follows: In the first few rounds of training for federated learning, collect the gradient updates uploaded by normal nodes ; Gradient updates uploaded to normal nodes , normalized to meet the requirements of the input neural network: ; in, represents the L2 norm of the gradient; is the normalized input; Using normal gradient as training data set, optimizing reconstruction error, the autoencoder contains encoder and decoder ;The optimization goal is to minimize the reconstruction error, and the loss function is MSE; In each round of federated learning training, the uploaded gradient is checked before the gradient is updated. Perform standardized preprocessing; Use the trained autoencoder to reconstruct the gradient and calculate the reconstruction error : ; Compared with the detection threshold δ: if >δ, the gradient is judged to be abnormal and may come from a malicious node; if ≤δ, the gradient is judged to be normal.

[0012] A system for desensitizing oral medical privacy data based on federated learning, comprising a processor, a memory, and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the method for desensitizing oral medical privacy data based on federated learning as described above.

[0013] The present invention has the following beneficial effects: 1. The present invention can achieve the coordinated optimization of privacy protection and diagnostic model performance, and provide a safe, efficient and sustainable solution for cross-regional and multi-institutional medical collaboration; 2. The present invention allows oral medical data to participate in collaborative training while being stored locally through federated learning, without the need for data to be discharged from the hospital, fundamentally reducing the risk of privacy leakage caused by data flow. The privacy parameters are dynamically adjusted through reinforcement learning, and the noise scale can be dynamically increased or decreased according to the current training round and model accuracy status, achieving a refined balance between privacy protection and model performance. 3. The present invention uses differential privacy technology, intelligent adjustment, and gradient anomaly detection mechanism for multiple encryption protections to effectively curb malicious attacks and ensure data security. It also uses high-quality data combined with dynamic privacy parameter adjustment to ensure efficient convergence of federated learning and improve the overall diagnostic capability of the model, prevent malicious updates, and significantly reduce the impact of unstable factors on global model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0015] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments: refer to Figure 1 In this embodiment, a method for desensitizing oral medical privacy data based on federated learning is provided, comprising the following steps: Acquire multi-source data related to oral healthcare, and classify, clean and desensitize different types of medical data; Based on the processed high-quality data, a federated learning system is built to achieve collaborative modeling across dental medical institutions; By intelligently and dynamically adjusting privacy parameters and attack defense, the security and model performance of the federated learning system are improved as follows: Introducing a reinforcement learning algorithm to dynamically adjust differential privacy parameters (such as noise intensity and privacy budget ε) and control the noise amplitude according to the current training round and model accuracy to minimize the impact on model performance; An anomaly detection mechanism is introduced in the transmission gradient to monitor whether the uploaded updates are abnormal and defend against data poisoning or gradient back-propagation attacks.

[0016] In this embodiment, multi-source data related to oral medical care is obtained, and different types of medical data are classified, cleaned and desensitized, as follows: The medical data includes oral imaging data (such as X-rays, 3D modeling, dental photos), structured data (patient medical records, dental scoring sheets, etc.) and unstructured text data (doctor diagnosis records, patient descriptions); Check and standardize data formats (such as image resolution and field content consistency), and remove redundant fields and useless data; Oral image data uses a deep learning segmentation algorithm to separate the tooth area and facial information, removing unnecessary privacy features; Structured data processing involves hashing or replacing sensitive fields (such as address and contact information); Unstructured text data uses NLP models (such as the BERT-based NER model) to automatically detect and mask sensitive information such as patient names and ID numbers.

[0017] In this embodiment, the oral image data uses a deep learning segmentation algorithm to separate the tooth area and facial information and remove unnecessary privacy features, as follows: Obtain oral imaging data, including dental X-rays and CT slices, and standardize the oral imaging data to a fixed resolution to ensure that the deep learning model supports consistent size: ; in, The height and width corresponding to the target resolution; is the original image; is the resized image; is a pixel; is the size adjustment function; According to the professional annotation tool of stomatology, the segmentation mask of the tooth area in the image is manually annotated to form the training label M(x,y); Build a segmentation model based on the UNet model, which consists of an encoder and a decoder and supports high-precision pixel-level segmentation; Encoder: progressively downsamples the image and extracts multi-scale features using multiple layers of convolution and pooling; Decoder: progressively upsamples, uses skip connections to pass encoder features to the decoder, and combines high- and low-level feature information; According to the training data set, input the image and the true label mask M(x,y), calculate the loss L through forward propagation, and use the Adam optimizer to update the model parameters; Input a new unlabeled oral image, and output the tooth area segmentation result after the segmentation model predicts it. , Based on the tooth region segmentation results , the non-tooth area is set to black or pseudo data noise to generate a desensitized image.

[0018] In this embodiment, a segmentation model is constructed based on the UNet model, as follows: In the encoder part, two convolution operations are performed consecutively: ; Among them, is the feature map, is the convolution kernel, is the bias, is a nonlinear activation function; For input data; Use Max-Pooling to gradually downsample the feature map: ; Among them, P(X) represents the output feature map after downsampling; K=k*l is the pooling window size; is the maximum value of all values ​​in the pooling window; X is the input feature map; In the decoder part, the spatial resolution of the feature map is increased by upsampling: ; Among them, U(X) represents the output feature map after upsampling; Indicates upsampling; scale=2 is the upsampling multiple, which means that the resolution of the feature map in each dimension is increased by 2 times; Each layer of decoder is concatenated with the corresponding encoder feature map through skip connection: ; Among them, C is the concatenated feature map; Concat represents the concatenation operation; are the outputs of the decoder and encoder respectively; Output predicted segmentation mask , whose value is a probability distribution value between [0,1]; During training, a weighted combination of BCE and Dice is used as the loss function L: ; ; ; Among them, λ 1 ,λ 2 represents the weight of BCE and Dice Loss; L BCE is the BCE loss (for pixel-level prediction accuracy); L Dice is the Dice loss (optimizing the segmentation performance of the target area), and N is the number of samples.

[0019] In this embodiment, the federated learning system includes independent nodes and a central server. Specifically: Treat each medical institution as an independent node, use its local cleaned data to participate in distributed training, upload only model parameters (such as weights or gradients) without sharing original data; the node local model processes specific tasks (such as tooth area segmentation, caries detection, case prediction, etc.); Each node Using its local dataset , according to the current weight W t Perform local gradient updates :

[0020] in, η is the learning rate; For Node Loss functions for medical tasks such as segmentation or classification; represents the gradient function; Central server: schedules training tasks and collects model updates from all nodes After that, perform global model aggregation (such as FedAvg algorithm) to generate the optimized global model weight W t+1 : ; Where n is the number of nodes; The optimized global model weights are distributed back to each node for the next round of iteration.

[0021] In this embodiment, a reinforcement learning algorithm is introduced to dynamically adjust the differential privacy parameters, and the noise amplitude is controlled according to the current training round and model accuracy to minimize the impact on model performance, as follows: a. Initialize the federated learning global model W 0 , Initial Privacy Budget , noise intensity σ 0 ; Start the reinforcement learning agent and generate privacy parameter adjustment actions according to the initial strategy; b. Each node Use the privacy budget at the current time t and the noise intensity σ t , in the local dataset The model training is completed on the GitHub repository; the implementation formula for differential privacy is: ; in, is the gradient update; Gaussian noise generated by the current noise intensity; For Node On-premises data The gradient of the loss function L on the model parameters; I is the identity matrix; represents differential privacy noise; c. The central server aggregates the updates uploaded by the nodes to obtain the global model; d. According to the current status , using reinforcement learning models to generate new privacy parameter adjustment actions ; Among them, L t is the loss function at the current moment; are the adjustment amounts of privacy budget and noise intensity respectively; Execute actions to update privacy parameters: ; Use RL algorithms to update reinforcement learning policy network weights; e. Privacy parameters Model weight W t+1 Distribute to nodes and enter the next round of training, and loop until the termination condition is met (such as reaching the target model accuracy or the maximum number of training rounds).

[0022] In this embodiment, an anomaly detection mechanism is introduced in the transmission gradient to monitor whether the uploaded update is abnormal and defend against data poisoning or gradient back-pushing attacks, as follows: In the first few rounds of federated learning training (or from historical gradient distribution information), collect gradient updates uploaded by normal nodes ; Gradient updates uploaded to normal nodes , normalized to meet the requirements of the input neural network: ; in, represents the L2 norm of the gradient; is the normalized input; Using normal gradient as training data set, optimizing reconstruction error, the autoencoder contains encoder and decoder ;The optimization goal is to minimize the reconstruction error, and the loss function is MSE; In each round of federated learning training, the uploaded gradient is checked before the gradient is updated. Perform standardized preprocessing; Use the trained autoencoder to reconstruct the gradient and calculate the reconstruction error : ; Compared with the detection threshold δ: if >δ, the gradient is judged to be abnormal and may come from a malicious node; if ≤δ, the gradient is judged to be normal.

[0023] A system for desensitizing oral medical privacy data based on federated learning, comprising a processor, a memory, and a computer program stored in the memory. When the processor executes the computer program, the steps in the method for desensitizing oral medical privacy data based on federated learning are specifically performed. It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0024] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0025] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0026] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0027] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. A method for desensitizing oral medical privacy data based on federated learning, characterized in that: The following steps are involved: Acquire multi-source data related to oral healthcare, and classify, clean and desensitize different types of medical data; Based on the processed high-quality data, a federated learning system is built to achieve collaborative modeling across dental medical institutions; By intelligently and dynamically adjusting privacy parameters and attack defense, the security and model performance of the federated learning system are improved as follows: Introducing a reinforcement learning algorithm to dynamically adjust differential privacy parameters and control the noise amplitude according to the current training round and model accuracy to minimize the impact on model performance; An anomaly detection mechanism is introduced in the transmission gradient to monitor whether the uploaded updates are abnormal and defend against data poisoning or gradient back-propagation attacks.

2. According to claim 1, a method for desensitizing oral medical privacy data based on federated learning is characterized in that: The method of obtaining multi-source data related to oral medical care and classifying, cleaning and desensitizing different types of medical data is as follows: The medical data includes oral imaging data, structured data and unstructured text data; Check and standardize data formats, remove redundant fields and useless data; Oral image data uses a deep learning segmentation algorithm to separate the tooth area and facial information, removing unnecessary privacy features; Structured data processing involves hashing or replacing sensitive fields; Unstructured text data uses NLP models to automatically detect and mask sensitive information such as patient names and ID numbers.

3. According to claim 2, a method for desensitizing oral medical privacy data based on federated learning is characterized in that: The oral image data uses a deep learning segmentation algorithm to separate the tooth area and facial information and remove unnecessary privacy features, as follows: Obtain oral imaging data, including dental X-rays and CT slices, and standardize the oral imaging data to a fixed resolution to ensure that the deep learning model supports consistent size: ; in, The height and width corresponding to the target resolution; is the original image; is the resized image; is a pixel; is the size adjustment function; According to the professional annotation tool of stomatology, the segmentation mask of the tooth area in the image is manually annotated to form the training label M(x,y); A segmentation model is built based on the UNet model, which consists of an encoder and a decoder and supports high-precision pixel-level segmentation; Encoder: progressively downsamples the image and extracts multi-scale features using multiple layers of convolution and pooling; Decoder: progressively upsamples, uses skip connections to pass encoder features to the decoder, and combines high- and low-level feature information; According to the training data set, input the image and the true label mask M(x,y), calculate the loss L through forward propagation, and use the Adam optimizer to update the model parameters; Input a new unlabeled oral image, and output the tooth area segmentation result after the segmentation model predicts it. , Based on the tooth region segmentation results , the non-tooth area is set to black or pseudo data noise to generate a desensitized image.

4. According to claim 3, a method for desensitizing oral medical privacy data based on federated learning is characterized in that: The segmentation model is constructed based on the UNet model, as follows: In the encoder part, two convolution operations are performed consecutively: ; Among them, is the feature map, is the convolution kernel, is the bias, is a nonlinear activation function; For input data; Use max pooling to progressively downsample the feature map: ; Among them, P(X) represents the output feature map after downsampling; K=k*l is the pooling window size; is the maximum value of all values ​​in the pooling window; X is the input feature map; In the decoder part, the spatial resolution of the feature map is increased by upsampling: ; Among them, U(X) represents the output feature map after upsampling; Indicates upsampling; scale=2 is the upsampling multiple, which means that the resolution of the feature map in each dimension is increased by 2 times; Each layer of decoder is concatenated with the corresponding encoder feature map through skip connection: ; Among them, C is the concatenated feature map; Concat represents the concatenation operation; are the outputs of the decoder and encoder respectively; Output predicted segmentation mask , the predicted segmentation mask value is a probability distribution value between [0,1]; During training, a weighted combination of BCE and Dice is used as the loss function L: ; ; ; Among them, λ1,λ2 represent the weights of BCE and Dice Loss; L BCE is the loss of BCE; L Dice is the Dice loss, and N is the number of samples.

5. According to the method for desensitizing oral medical privacy data based on federated learning in claim 1, it is characterized in that: The federated learning system includes independent nodes and a central server, specifically: Treat each medical institution as an independent node, use the local cleaned data of the medical institution to participate in distributed training, upload only model parameters without sharing original data; the node local model processes specific tasks; Each node Using its local dataset , according to the current weight W t Perform local gradient updates : ; in, η is the learning rate; For Node Loss function for medical tasks; represents the gradient function; Central server: schedules training tasks and collects model updates from all nodes After that, the global model is aggregated to generate the optimized global model weight W t+1 : ; Where n is the number of nodes; The optimized global model weights are distributed back to each node for the next round of iteration.

6. According to claim 5, a method for desensitizing oral medical privacy data based on federated learning is characterized in that: The reinforcement learning algorithm is introduced to dynamically adjust the differential privacy parameters and control the noise amplitude according to the current training round and model accuracy to minimize the impact on model performance. The details are as follows: a. Initialize the federated learning global model W0 and the initial privacy budget , noise intensity σ0; Start the reinforcement learning agent and generate privacy parameter adjustment actions according to the initial strategy; b. Each node Use the privacy budget at time t and the noise intensity σ t , in the local dataset The model training is completed on the GitHub repository; the implementation formula for differential privacy is: ; in, is the gradient update; Gaussian noise generated by the current noise intensity; For Node On-premises data The gradient of the loss function L on the model parameters; I is the identity matrix; represents differential privacy noise; c. The central server aggregates the updates uploaded by the nodes to obtain the global model; d. According to the current status , using reinforcement learning models to generate new privacy parameter adjustment actions ; Among them, L t is the loss function at the current moment; are the adjustment amounts of privacy budget and noise intensity respectively; Execute actions to update privacy parameters: ; Use RL algorithms to update reinforcement learning policy network weights; e. Privacy parameters Model weight W t+1 Distribute to nodes and enter the next round of training, and loop until the termination condition is met.

7. According to claim 5, a method for desensitizing oral medical privacy data based on federated learning is characterized in that: The anomaly detection mechanism is introduced in the transmission gradient to monitor whether the uploaded update is abnormal and defend against data poisoning or gradient back-pushing attacks. The details are as follows: In the first few rounds of training for federated learning, collect the gradient updates uploaded by normal nodes ; Gradient updates uploaded to normal nodes , normalized to meet the requirements of the input neural network: ; in, represents the L2 norm of the gradient; is the normalized input; Using normal gradient as training data set, optimizing reconstruction error, the autoencoder contains encoder and decoder ;The optimization goal is to minimize the reconstruction error, and the loss function is MSE; In each round of federated learning training, the uploaded gradient is checked before the gradient is updated. Perform standardized preprocessing; Use the trained autoencoder to reconstruct the gradient and calculate the reconstruction error : ; Compared with the detection threshold δ: if >δ, the gradient is judged to be abnormal and may come from a malicious node; if ≤δ, the gradient is judged to be normal.

8. An oral medical privacy data desensitization system based on federated learning, characterized in that: It includes a processor, a memory, and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the oral medical privacy data desensitization method based on federated learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Medical image automatic segmentation method based on deep learning

    CN114066866A

  • Drug administration method and device based on privacy protection and image recognition

    CN114519380A

  • Large model cutting federated learning method and system based on local features

    CN117521856A

  • Next-generation point-of-interest security recommendation strategy based on asynchronous advantage reinforcement learning

    CN118484596A

  • Data privacy protection method under federated learning framework and medical service system

    CN119312400A

Cited By

  • Health data enhancement and anomaly detection method based on GAN (Generative Adversarial Network)

    CN120277590A

  • Health data enhancement and anomaly detection method based on generative adversarial network (GAN)

    CN120277590B

  • Medical data privacy protection method based on federated learning and bert model

    CN120316805A

  • Multi-technology fused multi-source data grading and classifying method and system

    CN120429765A

  • A multi-technology integrated method and system for hierarchical classification of multi-source data

    CN120429765B