Collaborative optimization system and method based on differential privacy protection and model interpretability

By adopting a collaborative optimization system of differential privacy protection and model interpretability in the target identification classification model, dynamically adjusting the model's loss function and privacy constraints, the privacy protection and interpretability problems in the application of the model in high-risk fields is solved, and the balance of classification performance, privacy protection and interpretability is achieved.

CN120145441APending Publication Date: 2025-06-13ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510206824.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing target identification classification models are difficult to achieve effective privacy protection and model interpretability when processing sensitive data, resulting in limited applications in high-risk areas.

Method used

A collaborative optimization system based on differential privacy protection and model interpretability is adopted. Through the indicator integration module, dynamic balance module and verification evaluation module, the model's loss function, differential privacy constraints and interpretability indicators are dynamically adjusted to achieve a balance of classification performance, privacy protection strength and interpretability.

Benefits of technology

While maintaining high classification performance, it provides appropriate privacy protection and reasonable model explanations, achieving the optimal balance point in different application scenarios, and is suitable for medical diagnosis, financial risk control and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145441A_ABST
    Figure CN120145441A_ABST
Patent Text Reader

Abstract

The invention discloses a collaborative optimization system and method based on differential privacy protection and model interpretability, and belongs to the technical field of artificial intelligence. The system comprises an index integration module, a dynamic balance module and a verification evaluation module; the classification performance, the privacy protection intensity and the interpretability performance of the target identification classification model are evaluated and monitored based on an index integration module, and then a loss function, a privacy budget, a gradient cutting threshold and an interpretation method of the classification model are simultaneously adjusted based on a dynamic adjustment module; and the adjusted indexes are fed back to the system to find an optimal balance point, and multi-task collaborative optimization is completed. According to the method, a collaborative optimization system is constructed, the loss function, the differential privacy constraint and the interpretability index of the classification task are integrated into a unified optimization target, the classification performance, the privacy protection strength and the interpretability are optimized at the same time through a dynamic balance mechanism, and the optimal balance point of the three targets is found in different application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a collaborative optimization system and method based on differential privacy protection and model interpretability. Background Art

[0002] With the rapid development of artificial intelligence technology, target recognition and classification models have been widely used in many fields such as medical diagnosis, financial risk control, and autonomous driving. However, there are some problems in actual use: First, a large amount of sensitive data containing personal privacy information often needs to be processed during the model training process; Second, these models are often regarded as "black boxes", and data owners cannot see what actions or reasons the model uses to obtain the judgment results. Users often have doubts about the decisions made by the model, which limits its application in high-risk fields.

[0003] To solve this problem, various model interpretability methods have been proposed currently, mainly including the following categories: Gradient-based explanation methods, which explain model decisions by calculating the gradient of the model output with respect to the input; Perturbation-based methods, which observe the changes in the model output by systematically changing the input; Shapley value-based methods, which borrow concepts from game theory to quantify feature importance; Activation value-based methods, which analyze the activation patterns of the intermediate layers of neural networks; and Class activation-based methods, which focus on the activation regions of the model during specific class predictions. The effectiveness of these methods can be evaluated from three dimensions: The fidelity of model interpretation is used to measure the consistency between the interpretation and the model decision; Robustness is used to evaluate the stability of the model interpretation method to input perturbations; and Simplicity is used to measure the conciseness of the model interpretation.

[0004] At the same time, during the training process, the model often needs to process a large amount of sensitive data, which may contain personal privacy information. In recent years, with the introduction and improvement of data protection regulations in various countries, data privacy protection has become an important issue that cannot be ignored in the field of machine learning. Therefore, for the problem of privacy leakage of target recognition and classification models, certain privacy protection measures need to be taken to protect the privacy of data providers and promote the development of artificial intelligence technology.

[0005] Currently, in terms of the privacy protection of models, differential privacy has become one of the most promising privacy protection technologies. Differential privacy is a mathematical framework for achieving privacy protection in data distribution and query results. Its core is to introduce randomness to mask personal information in query results, ensuring that the output of the algorithm will not change drastically due to the addition or deletion of a single data record, thereby protecting the privacy of personal data. The specific protection method is to inject noise during the training process to prevent the leakage of personal information. In this way, attackers cannot steal the original data of the model through the designed attack model. The mainstream differential privacy protection methods include: differential privacy stochastic gradient descent (DPSGD) and its variant methods.

[0006] While enhancing privacy protection and interpretability, it is also necessary to ensure that the model maintains a high prediction accuracy. However, the noise introduced by the differential privacy protection mechanism will affect the model's ability to extract and recognize features, thereby affecting the interpretability and prediction accuracy of the model. Although existing research has been committed to combining interpretability and privacy, balancing the trade-off between data privacy protection and model performance, there is still a lack of systematic research and solutions on how to optimize these three goals of privacy protection, interpretability, and prediction accuracy simultaneously. Summary of the Invention

[0007] The purpose of the present invention is to provide a collaborative optimization system and method based on differential privacy protection and model interpretability, which is used to simultaneously optimize the classification performance, differential privacy protection, and interpretability of the target recognition and classification model through a dynamic balance mechanism.

[0008] To achieve the above invention purpose, a collaborative optimization system based on differential privacy protection and model interpretability provided by an embodiment includes an index integration module, a dynamic balance module, and a verification and evaluation module, which are used to simultaneously optimize the classification performance, differential privacy protection, and interpretability of the target recognition and classification model, and achieve the optimal balance point of the target recognition and classification model in different application scenarios;

[0009] The index integration module is used to integrate the loss function, differential privacy constraint, and interpretability index of the target recognition and classification model to evaluate the performance, privacy protection intensity, and interpretability performance of the target recognition and classification model;

[0010] The dynamic balance module is used to, based on the evaluation results of the classification performance, privacy protection intensity, and interpretability performance of the target recognition and classification model, simultaneously adjust the loss function and differential privacy constraint of the target recognition and classification model by combining an adaptive parameter adjustment strategy, balance the performance and privacy protection intensity of the target recognition and classification model, and combine the interpretability index to obtain the interpretation of the target recognition and classification model by dynamically evaluating and converting the selected interpretation method;

[0011] The verification and evaluation module is used to monitor the performance of the adjusted target recognition and classification model, the privacy protection intensity, and the interpretability performance, and feedback the loss function, differential privacy constraints, and interpretability metrics of the adjusted target recognition and classification model to the metric integration module and the dynamic balance module for simultaneously evaluating and optimizing the classification performance, differential privacy protection intensity, and interpretability performance of the target recognition and classification model.

[0012] In one embodiment, the classification performance is evaluated by the prediction accuracy, precision, and recall rate of the target recognition and classification model; the privacy protection intensity is evaluated by calculating the actual privacy guarantee level of the target recognition and classification model through differential privacy accounts; the interpretability performance is evaluated by the stability of the feature importance ranking based on the selected interpretation method.

[0013] In one embodiment, the loss function for adjusting the target recognition and classification model includes: measuring the difference between the model prediction result and the true label by minimizing the cross-entropy loss function, and based on the difference between the prediction result and the true label, using dropout and L2 regularization to optimize the parameters of the target recognition and classification model.

[0014] In one embodiment, the adjustment of differential privacy constraints includes: initializing the privacy budget and gradient clipping threshold based on the differential privacy stochastic gradient descent method, and then performing gradient update through clipping and noise injection to obtain the final gradient, thereby enhancing the privacy protection intensity.

[0015] In one embodiment, balancing the performance of the target recognition and classification model and the privacy protection intensity includes: adjusting the privacy budget and gradient clipping threshold within a relatively small range of variation, recording the performance changes of the target recognition and classification model after each adjustment, establishing a database of the performance and privacy protection intensity effects of the target recognition and classification model, and finding the balance point between the performance of the target recognition and classification model and the privacy protection intensity.

[0016] In one embodiment, the interpretability evaluation metrics include faithfulness, stability, and simplicity;

[0017] The faithfulness is measured by calculating the faithfulness loss function of the selected interpretation method, which is used to measure the consistency between the interpretation result of the classification model and the prediction behavior of the target recognition and classification model. The faithfulness loss function is denoted as L faithful , and the calculation formula is as follows:

[0018] L faithful =Σ|f(x)-f(mask k (x, φ(f, x)))|

[0019] Among them, x represents the input of the selected interpretation method, f(x) represents the interpretation of the target recognition classification model based on x, φ(f, x) represents the importance score of the output feature f with x as the input based on the selected interpretation method, and mask k (x, φ(f, x)) represents retaining the most important k features in the input x according to the importance score φ(f, x) of the output feature f, and setting the other features to zero;

[0020] The selected stability is calculated by the robustness loss function of the selected interpretation method to ensure the stability of the interpretation result to input perturbations. The robustness loss function is denoted as L robust , and the calculation formula is as follows:

[0021] L robust =Σ||φ(f, x)-φ(f′, x′)|| 2

[0022] Among them, x represents the input of the selected interpretation method, x′ represents the sample generated by adding perturbations to x, φ(f, x) is the importance score of the output feature f with x as the input based on the selected interpretation method, and φ(f, x′) is the importance score of the output feature f′ with the sample generated by adding perturbations to x as the input based on the selected interpretation method;

[0023] The selected simplicity is calculated by the simplicity loss function of the selected interpretation method to prompt the interpretation result to focus on a small number of key features. The simplicity loss function is denoted as L sparse , and the calculation formula is as follows:

[0024] L sparse =λ||φ(f, x)|| 1

[0025] Among them, λ is the regularization coefficient, and φ(f, x) is the importance score of the output feature f with x as the input based on the selected interpretation method.

[0026] In one embodiment, the method for obtaining the interpretation of the target recognition classification model by dynamically evaluating and converting the selected interpretation method includes: generating a benchmark interpretation of the target recognition classification model based on the feature importance method, calculating the interpretability index, dynamically selecting the interpretation method based on the interpretability index and the time overhead of interpretation generation, and obtaining the interpretation of the target recognition classification model.

[0027] In one embodiment, the dynamic selection of the explanation method based on the interpretability index and the time overhead of explanation generation includes: calculating the faithfulness loss function of the selected explanation method to judge the credibility of the explanation method; calculating the robustness loss function and comparing it with a threshold to judge the stability of the explanation method; calculating the time for explanation generation and comparing it with a preset time threshold; and switching to a suitable explanation method by judging the credibility, stability, and time for explanation generation.

[0028] The present invention also provides a collaborative optimization method based on differential privacy protection and model interpretability. The collaborative optimization method applies the collaborative optimization system based on differential privacy protection and model interpretability, and includes the following steps:

[0029] Collect the data required by the target recognition and classification model and preprocess the data as input, initialize the model parameters, privacy budget, gradient clipping threshold, and explanation method, and perform target recognition and classification model training;

[0030] During the training of the target recognition and classification model, evaluate and monitor the classification performance, privacy protection strength, and interpretability performance of the target recognition and classification model based on the metric integration module;

[0031] Based on the results of the evaluation and monitoring, adjust the loss function, privacy budget, gradient clipping threshold, and explanation method of the target recognition and classification model in the dynamic adjustment module to balance the model performance, privacy protection strength, and model interpretability;

[0032] Feed back the adjusted loss function, differential privacy constraint, and interpretability index to the metric integration module and the dynamic balance module for simultaneously evaluating and optimizing the classification performance, differential privacy protection, and interpretability of the target recognition and classification model to find the optimal balance point and complete the collaborative optimization of the classification performance, differential privacy protection, and interpretability of the target recognition and classification model.

[0033] Compared with the prior art, the beneficial effects of the present invention at least include:

[0034] (1) Integrate the classification performance, privacy protection strength, and interpretability performance into a unified optimization goal, and simultaneously optimize these three metrics through a dynamic balance mechanism, which can provide appropriate privacy protection and reasonable model interpretation while maintaining a high classification performance.

[0035] (2) By monitoring the classification performance, privacy protection strength, and interpretability performance metrics of the target recognition and classification model, and feeding back the results of the metrics to the metric integration module and the dynamic balance module to form a closed-loop optimization, the optimal balance point of these three goals can be found in different application scenarios.

[0036] (3) The generality of this method enables it to be applicable not only to medical diagnosis but also to be extended to other fields with high requirements for privacy protection and interpretability, such as financial risk control and autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art.

[0038] Figure 1 Schematic structural diagram of a collaborative optimization system based on differential privacy protection and model interpretability provided for the embodiment;

[0039] Figure 2 Schematic structural diagram of the dynamic balance module and the verification and evaluation module provided for the embodiment;

[0040] Figure 3 Schematic flowchart of a collaborative optimization method based on differential privacy protection and model interpretability provided for the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the following further details the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0042] To simultaneously optimize the classification performance, differential privacy protection strength and interpretability of the target recognition and classification model, the embodiment provides a collaborative optimization system based on differential privacy protection and model interpretability, as Figure 1 shown. The system includes an index integration module, a dynamic balance module and a verification and evaluation module. Each part of the system architecture will be described in detail below.

[0043] The index integration module is used to integrate the loss function, differential privacy constraint and interpretability index of the target recognition and classification model to evaluate the performance, privacy protection strength and interpretability performance of the target recognition and classification model.

[0044] In the embodiment, the cross-entropy loss function L classification is adopted for the target recognition and classification task, which is used to measure the difference between the prediction result of the target recognition and classification model and the true label. For the N-classification problem, the cross-entropy loss calculates the negative logarithmic likelihood probability of each sample on its true class and takes the average of all samples.

[0045]

[0046] where m is the number of samples input to the target recognition and classification model, N is the number of classes, and y ijDenote the true label that the i-th sample belongs to the j-th class, p ij is the probability that the target recognition and classification model predicts that the i-th sample belongs to the, class.

[0047] In the embodiment, on the basis of ensuring the target recognition and classification performance, a differential privacy protection mechanism is introduced. The perturbation of the gradient is realized through the Differentially Private Stochastic Gradient Descent (DPSGD) method to achieve the privacy data protection of the target recognition and classification model; wherein DPSGD involves two key parameters: the privacy budget ε and the gradient clipping threshold C. Each time the gradient is updated, first limit the L2 norm of the sample gradient within the threshold C, and then add Gaussian noise that conforms to the N(0, σ 2 ) distribution, where:

[0048]

[0049] In the embodiment, in terms of the interpretability of the target recognition and classification model, the system dynamically selects and switches different interpretation methods, where the interpretation methods include: the gradient-based interpretation method InputXGradient and the feature-importance-based interpretation methods such as SHAP and LIME, and the performance of the model interpretability is evaluated through three dimensions: faithfulness, stability, and simplicity.

[0050] Among them, the faithfulness loss function: measures the consistency between the interpretation result and the model prediction behavior. Let φ(f, x) be the importance score of the output features of the selected interpretation function (at least one of SHAP and InputXGradient). For any input x, if the k most important features are found according to φ(f, x) and the values outside these features are set to zero to obtain x′, then f(x) and f(x′) should be as close as possible. Therefore, the faithfulness loss function is defined as:

[0051] L faithful =∑|f(x)-f(mask k (x, φ(f, x)))|

[0052] Among them, x represents the input of the selected interpretation method, f(x) represents the interpretation of the target recognition and classification model based on x, φ(f, x) is the importance score of the output feature f of the selected interpretation method with x as the input, and mask k (x, φ(f, x)) represents retaining the k most important features in the input x according to the importance score φ(f, x) of the output feature f and setting other features to zero. This loss function ensures that the important features identified by the interpretation method do play a key role in the model prediction.

[0053] Robustness loss function: Ensure that the interpretation results of the target recognition and classification model are stable against input perturbations. Add perturbations to the input sample x of the selected interpretation method to generate x′, and require the original interpretation to be consistent with the interpretation after perturbation. Therefore, the robustness loss function is defined as:

[0054] L robust =Σ||φ(f, x) - φ(f′, x′)|1 2

[0055] where x represents the input of the selected interpretation method, x′ represents the sample generated by adding perturbations to x, φ(f, x) is the importance score of the output feature f based on the selected interpretation method with x as the input, and φ(f, x′) is the importance score of the output feature f′ based on the selected interpretation method with the sample generated by adding perturbations to x as the input.

[0056] Simplicity penalty term: Promote the interpretation results to focus on a small number of key features. Achieved by applying L1 regularization to the importance scores of the features output by the interpretation function:

[0057] L sparse =λ||φ(f, x)|| 1

[0058] where λ is the regularization coefficient, and φ(f, x) is the importance score of the output feature f based on the selected interpretation method with x as the input. This makes the target recognition and classification model tend to highlight a small number of the most discriminative features during interpretation, improving the comprehensibility of the model interpretation.

[0059] Dynamic balance module: Used to balance the performance and privacy protection strength of the target recognition and classification model by combining an adaptive parameter adjustment strategy to simultaneously adjust the loss function and differential privacy constraints of the target recognition and classification model based on the evaluation results of the classification performance, privacy protection strength, and interpretability performance of the target recognition and classification model, and combine the interpretability metrics to obtain the interpretation of the target recognition and classification model by dynamically evaluating and transforming the selected interpretation method.

[0060] Verification and evaluation module: Used to monitor the adjusted classification performance, privacy protection strength, and interpretability performance, and feedback the adjusted loss function, differential privacy constraints, and interpretability metrics to the metric integration module and the dynamic balance module for simultaneously evaluating and optimizing the classification performance, differential privacy protection strength, and interpretability performance of the classification model.

[0061] Such as Figure 2The figure shows a schematic structural diagram of the dynamic balance module and the verification and evaluation module. In the embodiment, the parameter optimization process of the target recognition and classification model is directly guided by the loss function, and the optimization is carried out by minimizing the cross-entropy loss. To prevent overfitting, random inactivation with a dropout rate of 0.5 and L2 regularization are adopted, and the prediction accuracy, precision, and recall rate of the target recognition and classification model are monitored to optimize the classification performance.

[0062] To balance the classification performance and privacy protection intensity of the model target recognition and classification model, the system adopts an adaptive parameter adjustment strategy. Specifically, when a privacy risk is detected (such as the success rate of the membership inference attack exceeding the threshold), the privacy budget ε and the gradient clipping threshold C are reduced to enhance the privacy protection intensity; when the model performance drops significantly, the privacy budget ε and the gradient clipping threshold C of these parameters are appropriately relaxed. Specifically, the value range of ε is restricted within [1, 8], and the value range of C is restricted within [1, 10]. The step sizes of each adjustment are 0.5 and 0.2 respectively, and the actual privacy protection level is calculated through the differential privacy account to supervise and achieve the balance of the privacy protection intensity.

[0063] In the embodiment, the selection of the interpretation method is dynamically adjusted in combination with the interpretability index. When L faithful is relatively high, it indicates that the credibility of the current interpretation method is insufficient, and the system will preferentially switch to the SHAP method; when the interpretation result is too sensitive to slight changes in the input (L robust is lower than the threshold), it means that the interpretation result is not stable enough, and the system will also select a method with better theoretical guarantee. At the same time, the system will monitor the time overhead of generating the interpretation and switch to a lighter-weight method InputXGradient when the computing resources are limited (the time for generating the interpretation exceeds the preset threshold).

[0064] Based on this adaptive system parameter collaborative optimization mechanism, the system can provide appropriate privacy protection and reliable model interpretation while maintaining high classification performance.

[0065] Next, the system continuously monitors the effectiveness of the privacy protection of the target recognition and classification model, records the classification performance changes after each adjustment of the privacy budget and the gradient clipping threshold, and establishes an effect database for parameter adjustment to provide a reference for subsequent optimization. To comprehensively evaluate the effect of the system, three aspects of indicators are monitored on the validation set: classification performance (prediction accuracy, precision, recall rate), privacy protection intensity (the actual privacy guarantee level calculated through the differential privacy account), and interpretability performance (the stability of the feature importance ranking). The calculation results of these indicators are directly fed back to the indicator integration module and the dynamic balance module to form a closed-loop optimization.

[0066] Such as Figure 3As shown, the embodiment also provides a collaborative optimization method based on differential privacy protection and model interpretability, which is applied to the above collaborative optimization system and includes the following steps:

[0067] Collect the data required for the target recognition and classification model and preprocess the data as input. Initialize the model parameters, privacy budget, gradient clipping threshold, and interpretation method, and conduct target recognition and classification model training;

[0068] During the training of the target recognition and classification model, evaluate and monitor the classification performance, privacy protection intensity, and interpretability performance of the target recognition and classification model based on the metric integration module;

[0069] Based on the results of the evaluation and monitoring, adjust the loss function, privacy budget, gradient clipping threshold, and interpretation method of the target recognition and classification model in the dynamic adjustment module to balance the model performance, privacy protection intensity, and model interpretation;

[0070] Feed back the adjusted loss function, differential privacy constraint, and interpretability metric to the metric integration module and the dynamic balance module for simultaneously evaluating and optimizing the classification performance, differential privacy protection, and interpretability of the target recognition and classification model to find the optimal balance point and complete the collaborative optimization of the classification performance, differential privacy protection, and interpretability of the target recognition and classification model.

[0071] Next, taking the medical diagnosis scenario as an example, the specific implementation process of the collaborative optimization method will be described in detail.

[0072] In the medical diagnosis scenario, the input data of the target recognition and classification model includes various examination indicators of the patient (such as blood pressure, heart rate, blood sugar, etc.), and the goal is to predict whether the patient has a specific disease. This type of data has typical privacy sensitivity, and at the same time, doctors need to understand the diagnostic basis of the target recognition and classification model. Therefore, it is very suitable to apply the optimization method of the present invention.

[0073] In terms of model training, a three-layer fully connected neural network is selected as the basic classifier, and the number of neurons in each layer is 128, 64, and 2 respectively. The output layer corresponds to two states of healthy / sick. The target recognition and classification model parameters are optimized by minimizing the cross-entropy loss, and each mini-batch contains 32 samples. To prevent overfitting, random inactivation with a dropout rate of 0.5 and L2 regularization (coefficient λ = 0.01) are adopted.

[0074] When implementing differential privacy protection, the framework first initializes the key parameters: privacy budget ε = 4.0, gradient clipping threshold C = 3.0. Each gradient update goes through two steps: clipping and noise injection. Specifically, for the gradient g of each sample, first perform clipping:

[0075]

[0076] Among them, g clip represents the clipped gradient.

[0077] Then Gaussian noise is added to obtain the final gradient for updating:

[0078]

[0079] Among them, g noisy represents the updated gradient, represents the added Gaussian noise, and I represents random noise.

[0080] The system continuously monitors the effectiveness of privacy protection. When potential privacy risks are detected (e.g., through membership inference attack tests), the parameter adjustment mechanism is automatically triggered. The adjustment strategy adopts a progressive approach: each time the privacy budget ε or the gradient clipping threshold C is adjusted, the change range is limited within 0.5 and 0.2 respectively, to avoid sudden changes in model performance caused by drastic fluctuations. At the same time, the system records the performance changes of the target recognition and classification model after each adjustment, establishes an effect database for parameter adjustment, and provides a reference for subsequent optimization.

[0081] In terms of interpretability, the system implements a set of dynamic evaluation and switching mechanisms. First, the SHAP method is used to generate a baseline explanation, and the faithfulness loss function and robustness loss function are calculated to determine the faithfulness score and stability score. When the model predicts that a certain patient has a disease, the system will generate the key medical indicators affecting the prediction and their importance scores. If the time taken to generate the explanation exceeds the preset threshold (e.g., 0.5 seconds), the system will switch to the InputXGradient method with higher computational efficiency; if it is found that the explanation result is too sensitive to slight changes in the input (the robustness loss function is lower than the threshold), it will switch back to the SHAP method to obtain a more reliable explanation.

[0082] Finally, three aspects of indicators are monitored simultaneously on the validation set: classification performance (prediction accuracy, precision, recall), privacy protection strength (calculating the actual privacy guarantee level through differential privacy accounting), and interpretability performance (stability of feature importance ranking). The calculation results of these indicators are directly fed back to the parameter adjustment module to form a closed-loop optimization.

[0083] The experimental results show that in the diabetes diagnosis task, the collaborative optimization system can achieve the following performance: the prediction accuracy is maintained above 92%, while providing (4, 10 -5 )-differential privacy guarantee, and providing a stable and medically intuitive explanation for each diagnosis result (average stability score 0.85). Importantly, the optimization of these three goals is carried out collaboratively, avoiding the performance degradation that may occur when optimizing a single goal alone.

[0084] In this way, the present invention provides a practical solution for balancing the classification performance, privacy protection, and interpretability of the target recognition and classification model in practical applications. The generality of this method makes it applicable not only to medical diagnosis but also to other fields with high requirements for privacy protection and interpretability, such as financial risk control and autonomous driving.

[0085] The above-described specific embodiments have elaborated in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent substitutions, etc. made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A collaborative optimization system based on differential privacy protection and model interpretability, characterized in that: The collaborative optimization system includes an indicator integration module, a dynamic balance module and a verification and evaluation module, which are used to simultaneously optimize the classification performance, differential privacy protection and interpretability of the target recognition and classification model, so as to achieve the target recognition and classification model to find the optimal balance point in different application scenarios; The indicator integration module is used to integrate the loss function, differential privacy constraints and interpretability indicators of the target recognition and classification model to evaluate the performance, privacy protection strength and interpretability of the target recognition and classification model; The dynamic balancing module is used to balance the performance and privacy protection strength of the target recognition classification model based on the evaluation results of the classification performance, privacy protection strength and interpretability performance of the target recognition classification model by combining an adaptive parameter adjustment strategy and adjusting the loss function and differential privacy constraints of the target recognition classification model. In addition, the dynamic balancing module combines the interpretability index to dynamically evaluate and convert the selected interpretation method to obtain the interpretation of the target recognition classification model. The verification and evaluation module is used to monitor the performance, privacy protection strength and interpretability of the adjusted target recognition and classification model, and feed back the adjusted target recognition and classification model loss function, differential privacy constraints and interpretability indicators to the indicator integration module and the dynamic balance module, so as to simultaneously evaluate and optimize the classification performance, differential privacy protection strength and interpretability performance of the target recognition and classification model.

2. The collaborative optimization system according to claim 1, characterized in that: The classification performance is evaluated by the prediction accuracy, precision and recall of the target recognition classification model; the privacy protection strength is evaluated by calculating the actual privacy guarantee level of the target recognition classification model through a differential privacy account; the interpretability performance is evaluated by the stability of the feature importance ranking based on the selected explanation method.

3. The collaborative optimization system according to claim 1, characterized in that: The adjustment of the loss function of the target recognition classification model includes: measuring the difference between the model prediction result and the true label by minimizing the cross entropy loss function, and based on the difference between the prediction result and the true label, using random inactivation and L2 regularization to optimize the target recognition classification model parameters.

4. The collaborative optimization system according to claim 1, characterized in that: The adjusting of the differential privacy constraint includes: initializing the privacy budget and the gradient clipping threshold based on the differential privacy stochastic gradient descent method, and then updating the gradient by clipping and noise injection to obtain the final gradient, thereby enhancing the privacy protection strength.

5. The collaborative optimization system according to claim 1, characterized in that: The described balanced target recognition classification model performance and privacy protection strength includes: adjusting the privacy budget and gradient clipping threshold within a smaller range of change, and recording the performance changes of the target recognition classification model after each adjustment, establishing a target recognition classification model performance and privacy protection strength effect database, and finding the balance point between the target recognition classification model performance and privacy protection strength.

6. The collaborative optimization system according to claim 1, characterized in that: The interpretability evaluation indicators described include faithfulness, stability, and simplicity; The fidelity is calculated by calculating the fidelity loss function of the selected explanation method to measure the consistency between the explanation result of the classification model and the prediction behavior of the target recognition classification model. The fidelity loss function is expressed as L faithful , the calculation formula is as follows: L faithful =Σ|f(x)-f(mask k (x,φ(f,x)))| Where x represents the input of the selected explanation method, f(x) represents the object recognition classification model explanation based on x, φ(f, x) is the importance score of the output feature f based on the selected explanation method with x as input, and mask k (x, φ(f, x)) means retaining the most important k features in the input x according to the importance score φ(f, x) of the output feature f, and setting the other features to zero; The selected stability is used to ensure that the interpretation result is stable to the input perturbation by calculating the robustness loss function of the selected interpretation method. The robustness loss function is expressed as L robust , the calculation formula is as follows: L robust =Σ||φ(f,x)-φ(f′,x′)||2 Where x represents the input of the selected explanation method, x′ represents the sample generated by adding perturbations to x, φ(f, x) is the importance score of the output feature f based on the selected explanation method with x as input, and φ(f, x′) is the importance score of the output feature f′ based on the sample generated by adding perturbations to x as input of the selected explanation method; The selected simplicity is used to force the explanation results to focus on a small number of key features by calculating the simplicity loss function of the selected explanation method. The simplicity loss function is expressed as L sparse , the calculation formula is as follows: L sparse =λ||φ(f,x)||1 Where λ is the regularization coefficient and φ(f, x) is the importance score of the output feature f based on the selected explanation method with x as input.

7. The collaborative optimization system according to claim 1, characterized in that: The method of obtaining the explanation of the target recognition classification model by dynamically evaluating and converting the selected explanation method includes: generating a baseline explanation of the target recognition classification model based on a feature importance method, calculating an explainability index, and dynamically selecting an explanation method based on the explainability index and the time cost of explanation generation to obtain the explanation of the target recognition classification model.

8. The collaborative optimization system according to claim 7, characterized in that: The method of dynamically selecting an explanation method based on the explainability index and the time cost of explanation generation includes: calculating the fidelity loss function of the selected explanation method to judge the credibility of the explanation method; calculating the robustness loss function and comparing it with the threshold to judge the stability of the explanation method; calculating the time of explanation generation and comparing it with the preset time threshold; switching the appropriate explanation method by judging the credibility, stability and time of explanation generation.

9. A collaborative optimization method based on differential privacy protection and model interpretability, characterized in that: The collaborative optimization method applies the collaborative optimization system based on differential privacy protection and model interpretability according to any one of claims 1 to 8, comprising the following steps: Collect the data required for the target recognition and classification model and preprocess the data as input, initialize the model parameters, privacy budget, gradient clipping threshold and explanation method, and train the target recognition and classification model; In the training of the target recognition and classification model, the classification performance, privacy protection strength and interpretability of the target recognition and classification model are evaluated and monitored based on the indicator integration module; Based on the results of evaluation and monitoring, the loss function, privacy budget, gradient clipping threshold, and explanation method of the target recognition classification model are adjusted in the dynamic adjustment module to balance model performance, privacy protection strength, and model explanation. The adjusted loss function, differential privacy constraints and interpretability indicators are fed back to the indicator integration module and the dynamic balance module to simultaneously evaluate and optimize the classification performance, differential privacy protection and interpretability of the target recognition classification model, so as to find the optimal balance point and complete the collaborative optimization of the classification performance, differential privacy protection and interpretability of the target recognition classification model.