Green financial data differential privacy budget allocation method based on reinforcement learning
Through a reinforcement learning-based method, the privacy budget of green financial data is dynamically allocated, which solves the problems of inefficient static allocation and insufficient scenario adaptability in traditional methods, and realizes efficient, transparent and compliant privacy protection of green financial data.
Patent Information
- Application Number
- CN202510908939.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional differential privacy budget allocation methods have problems in green finance data, such as low static allocation efficiency, insufficient scenario adaptability, and lack of algorithm dynamics, which leads to an imbalance between privacy protection intensity and data value, and makes it difficult to cope with the dynamically changing carbon emission data and environmental benefit indicators in green finance scenarios.
A reinforcement learning-based method is adopted to dynamically allocate the privacy budget of green financial data through format-preserving encryption, generative adversarial networks, environmental benefit-privacy association matrix and target reinforcement learning model. Combined with multi-objective reward function and adaptive noise injection mechanism, Pareto optimal privacy budget allocation is achieved.
It achieves dynamic adaptability, scenario-based protection, and resource and energy efficiency optimization of the green finance data privacy budget, reduces privacy budget allocation errors, improves model accuracy and transparency, and meets financial regulatory requirements.
Smart Images

Figure CN120744974A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data processing, and in particular to a green financial data differential privacy budget allocation method based on reinforcement learning. Background Art
[0002] Green finance refers to economic activities that support environmental improvement, climate change response, and resource conservation and efficient utilization, namely, financial services provided for project investment and financing, project operations, risk management, etc. in the fields of environmental protection, energy conservation, clean energy, green transportation, green buildings, etc. Accordingly, green finance data can be understood as data related to green financial services, such as carbon emission data, production and operation data, third-party carbon verification reports, etc.
[0003] In order to ensure the security of green finance data and user privacy, existing technologies use differential privacy budget allocation methods to achieve the privacy security of green finance data. Specifically, they are used to dynamically allocate limited privacy protection resources (i.e., privacy budgets) during the release or query process of green finance data to balance the strength of privacy protection and data availability. However, traditional differential privacy budget allocation methods have multiple shortcomings: (1) Inefficiency of static allocation: Traditional methods use fixed budget allocation strategies (such as uniform allocation or empirical rules) and cannot adapt to the dynamic changes in carbon emission data, environmental benefit indicators and other high-dimensional features in green finance scenarios, resulting in an imbalance between privacy protection strength and data value; (2) Insufficient scenario adaptability: Traditional methods focus on general financial scenarios (such as credit scoring) and lack targeted protection mechanisms for multi-source heterogeneous data unique to green finance (such as carbon trading records and green project evaluation reports), resulting in the coexistence of privacy leakage risks and model accuracy loss; (3) Lack of algorithm dynamics: Traditional methods rely on manually set adjustment rules and are difficult to deal with the real-time changing privacy attack risks and data correlation in green finance data.
[0004] Therefore, how to effectively solve the problem of unreasonable privacy budget allocation for green financial data has become a research hotspot in this field. Summary of the Invention
[0005] This application provides a green finance data differential privacy budget allocation method based on reinforcement learning, aiming to allocate an appropriate privacy budget for green finance data.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] A reinforcement learning-based differential privacy budget allocation method for green financial data includes:
[0008] Perform format-preserving encryption on green finance data collected in real time to obtain desensitized financial data;
[0009] Inputting the desensitized financial data into a pre-trained generative adversarial network to obtain time series data with multiple features; the features are obtained by feature engineering based on historical green finance data;
[0010] Based on the time series data of the multiple features, combined with the correlation matrix method, an environmental benefit-privacy correlation matrix is constructed; the environmental benefit-privacy correlation matrix is used to quantify the contribution of the multiple features to the privacy leakage risk;
[0011] Based on the contribution of multiple features and time series data, they are input into a pre-built target reinforcement learning model to obtain the privacy budget allocation result output by the target reinforcement learning model; the target reinforcement learning model uses a multi-objective reward function to dynamically adjust the weight coefficient to achieve Pareto optimality, and the target reinforcement learning model also uses an adaptive noise injection mechanism to achieve self-matching of noise intensity and data fluctuation period based on the reinforcement learning action output; the adaptive noise injection mechanism includes a Laplace noise dynamic scaler; the privacy budget allocation result includes the privacy budget of the green financial data.
[0012] Optionally, the method further includes:
[0013] Obtaining local parameter information uploaded by each distributed node according to a federated training mechanism; the federated training mechanism is used to use the target reinforcement learning model as a federated model and train the federated model using local data to obtain model parameters as the local parameter information;
[0014] Utilize the local parameter information to perform parameter tuning on the target reinforcement learning model to achieve optimization of the target reinforcement learning model.
[0015] Optionally, the method further includes:
[0016] Based on the privacy budget of the green financial data, corresponding privacy-preserving data is generated, and the privacy-preserving data is published to a plurality of distributed nodes; the distributed nodes are pre-deployed with a trusted execution environment;
[0017] Based on a pre-deployed monitoring system, feedback data generated by each of the distributed nodes corresponding to the privacy-protected data is obtained, and each of the feedback data is used as new training data to train the target reinforcement learning model to achieve an update of the target reinforcement learning model.
[0018] Optionally, the method further includes:
[0019] Determining a SHAP value for each of the features during the learning process of the target reinforcement learning model; the SHAP value is used to quantify the contribution of the feature to the privacy budget allocation result;
[0020] Based on the SHAP value of each of the features, a corresponding explainability report is generated; the explainability report is used to meet the transparency requirements of financial supervision.
[0021] Optionally, the target reinforcement learning model includes a deep deterministic policy gradient model, which is used to model privacy budget allocation as a continuous action space optimization problem.
[0022] Optionally, the multi-objective reward function includes a triple reward, which includes a privacy protection reward, a data utility reward, and a resource cost reward. The privacy protection reward is negatively correlated with the probability of privacy leakage, the data utility reward is positively correlated with the accuracy of time series data, and the resource cost reward is negatively correlated with the consumption of computing resources.
[0023] Optionally, the Laplace noise dynamic scaler is based on long short-term memory network training.
[0024] A green financial data differential privacy budget allocation device based on reinforcement learning, comprising:
[0025] A data collection unit, used to encrypt the green financial data collected in real time in a format-preserving manner to obtain desensitized financial data;
[0026] A data processing unit, configured to input the desensitized financial data into a pre-trained generative adversarial network to obtain time series data of multiple features; the features are obtained by feature engineering based on historical green finance data;
[0027] A data association unit is configured to construct an environmental benefit-privacy association matrix based on the time series data of the plurality of features in combination with an association matrix method; the environmental benefit-privacy association matrix is configured to quantify the contribution of the plurality of features to the risk of privacy leakage;
[0028] A reinforcement learning unit is used to input the contribution of multiple features and time series data into a pre-built target reinforcement learning model to obtain a privacy budget allocation result output by the target reinforcement learning model; the target reinforcement learning model adopts a multi-objective reward function to dynamically adjust the weight coefficient to achieve Pareto optimality, and the target reinforcement learning model also adopts an adaptive noise injection mechanism to achieve self-matching of noise intensity and data fluctuation period based on the reinforcement learning action output; the adaptive noise injection mechanism includes a Laplace noise dynamic scaler; the privacy budget allocation result includes the privacy budget of the green financial data.
[0029] A storage medium comprising a stored program, wherein when the program is run by a processor, the method for allocating differential privacy budgets for green financial data based on reinforcement learning is executed.
[0030] An electronic device comprises: a processor, a memory and a bus; the processor and the memory are connected via the bus;
[0031] The memory is used to store programs, and the processor is used to run programs, wherein the program is executed by the processor to execute the reinforcement learning-based green financial data differential privacy budget allocation method.
[0032] The technical solution provided in this application performs format-preserving encryption on green financial data collected in real time to obtain desensitized financial data. The desensitized financial data is input into a pre-trained generative adversarial network to obtain time series data of multiple features. Based on the time series data of multiple features, combined with the association matrix method, an environmental benefit-privacy association matrix is constructed. Based on the contribution of multiple features and time series data, it is input into a pre-built target reinforcement learning model to obtain the privacy budget allocation result output by the target reinforcement learning model. This application uses the environmental benefit-privacy association matrix and the target reinforcement learning model to achieve reasonable allocation of the privacy budget of green financial data, which is objective and efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0034] Figure 1 A flowchart of a method for allocating differential privacy budgets for green financial data based on reinforcement learning provided in an embodiment of the present application;
[0035] Figure 2 A flowchart of another method for allocating differential privacy budgets for green financial data based on reinforcement learning provided in an embodiment of the present application;
[0036] Figure 3 A flowchart of another method for allocating differential privacy budgets for green financial data based on reinforcement learning provided in an embodiment of the present application;
[0037] Figure 4 A flowchart of another method for allocating differential privacy budgets for green financial data based on reinforcement learning provided in an embodiment of the present application;
[0038] Figure 5 A flowchart of another method for allocating differential privacy budgets for green financial data based on reinforcement learning provided in an embodiment of the present application;
[0039] Figure 6 A schematic diagram of the architecture of a green financial data differential privacy budget allocation device based on reinforcement learning provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0041] In this application, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or apparatus comprising the element.
[0042] like Figure 1 As shown, it is a flow chart of a green financial data differential privacy budget allocation method based on reinforcement learning provided in an embodiment of the present application, which includes the following steps.
[0043] S101: Perform format-preserving encryption on the green financial data collected in real time to obtain desensitized financial data.
[0044] Among them, the format preserving encryption (FPE) algorithm can be used to perform format-preserving encryption on the green financial data collected in real time to obtain desensitized financial data and improve data security.
[0045] S102: Input the desensitized financial data into a pre-trained generative adversarial network to obtain time series data of multiple features.
[0046] Among them, the features are obtained through feature engineering based on historical green finance data.
[0047] In some examples, a generative adversarial network can be pre-trained using historical samples, where the historical samples can be historical green finance data and corresponding feature time series data, where the feature time series data includes time series data of multiple features, such as carbon trading time series data.
[0048] S103: Based on the time series data of multiple features and combined with the correlation matrix method, an environmental benefit-privacy correlation matrix is constructed.
[0049] Among them, the environmental benefit-privacy association matrix is used to quantify the contribution of multiple features to the risk of privacy leakage.
[0050] In some examples, by constructing an environmental benefit-privacy association matrix, the contribution of multiple features such as carbon emission intensity and green project rating to the risk of privacy leakage is quantified and used as the core input of the state space of the target reinforcement learning model.
[0051] S104: Based on the contribution of multiple features and time series data, input them into a pre-built target reinforcement learning model to obtain the privacy budget allocation result output by the target reinforcement learning model.
[0052] Among them, the target reinforcement learning model adopts a multi-objective reward function to dynamically adjust the weight coefficient to achieve Pareto optimality, and the target reinforcement learning model also adopts an adaptive noise injection mechanism to achieve self-matching of noise intensity and data fluctuation period based on the reinforcement learning action output; the adaptive noise injection mechanism includes a Laplace noise dynamic scaler; the privacy budget allocation result includes the privacy budget of green financial data.
[0053] Optionally, the target reinforcement learning model includes a Deep Deterministic Policy Gradient (DDPG) model, which is used to model privacy budget allocation as a continuous action space optimization problem.
[0054] In some examples, the state space (i.e., continuous action space) designed by the DDPG model includes at least multiple features, and the multiple features include at least data sensitivity classification, real-time query frequency, environmental benefit indicator volatility, green project cycle, carbon emission intensity, differential privacy parameters, and query sensitivity.
[0055] In some examples, the target reinforcement learning model can be a dual-channel DDPG model. The dual channels include an environmental channel and a privacy channel. The environmental channel inputs include features related to environmental benefits (e.g., green project lifecycle, carbon emission intensity, etc.), while the privacy channel inputs include features related to privacy budget allocation (e.g., differential privacy parameters, query sensitivity, etc.). Furthermore, the output layer of the DDPG model is used to output the privacy budget allocation result. Specifically, the output layer jointly decides the privacy budget allocation vector to generate the privacy budget allocation result.
[0056] Optionally, the multi-objective reward function includes a triple reward, which includes a privacy protection reward, a data utility reward, and a resource cost reward. The privacy protection reward is negatively correlated with the probability of privacy leakage, the data utility reward is positively correlated with the accuracy of time series data, and the resource cost reward is negatively correlated with the consumption of computing resources.
[0057] In some examples, a triple reward mechanism is introduced into the target reinforcement learning model to dynamically adjust the weight coefficient (a hyperparameter) of the target reinforcement learning model to achieve Pareto optimality. Specifically, the key to achieving Pareto optimality through dynamic weight adjustment lies in real-time adjustment of the weight distribution of various features in the state space of the target reinforcement learning model based on environmental changes or decision-making requirements, thereby finding the optimal balance between the conflicting objectives corresponding to multiple features.
[0058] Optionally, the Laplace noise dynamic scaler is based on the Long Short-Term Memory (LSTM) network training.
[0059] In some examples, the adaptive noise injection mechanism combines the characteristics of time series data to design an LSTM-driven Laplace noise dynamic scaler. The LSTM input is time series data, and the output is Laplace noise.
[0060] Optionally, the performance of the target reinforcement learning model is closely related to the rationality of privacy budget allocation. In order to improve the rationality of privacy budget allocation, it is necessary to further improve the performance of the target reinforcement learning model. For specific methods to improve the performance of the target reinforcement learning model, please refer to Figure 2 and Figure 3 shown.
[0061] Optionally, in the field of financial supervision, the supervision of green financial data during use is relatively strict. Therefore, it is also necessary to perform feature attribution visualization on the learning process of the target reinforcement learning model. The implementation process of feature attribution visualization can be found in Figure 4 The method shown.
[0062] In some examples, combining 2- Figure 4 The method shown in the embodiment of the present application can be simply summarized as follows Figure 5 As shown in the figure, specifically, green financial data is collected first, followed by feature extraction and classification, then the environmental benefit-privacy association matrix is constructed, and then the reinforcement learning decision maker (i.e., the target reinforcement learning model) and adaptive noise injection are run, followed by privacy-protected data release, followed by federated model update, and finally multi-dimensional monitoring feedback is performed to update the reinforcement learning decision maker.
[0063] Compared with traditional methods, the embodiments of the present application establish a reinforcement learning strategy framework (i.e., a target reinforcement learning model) and design a multi-objective reward function based on the reinforcement learning strategy framework. The multi-objective reward function innovatively introduces a triple reward mechanism, and also constructs an environmental benefit-privacy association matrix to embed green features (i.e., contribution) into the reinforcement learning strategy framework, and injects an adaptive noise mechanism into the reinforcement learning strategy framework to achieve the update of the reinforcement learning strategy framework.
[0064] It should be noted that the advantages of the method shown in the embodiment of the present application are: (1) Breakthrough in dynamic adaptability: Compared with the adjustment of association attribute rules involved in other solutions, millisecond-level dynamic response is achieved through reinforcement learning, and the privacy budget allocation error is reduced by 63.2%; (2) Scenario-based technology innovation: For the first time, environmental benefit indicators are incorporated into the privacy decision-making system, and the accuracy of privacy budget allocation is improved by 41% in the carbon footprint tracking scenario; (3) Energy efficiency collaborative optimization: Through the computing resource consumption feedback mechanism, resource energy consumption is reduced by 27% compared with traditional methods at the same privacy strength; (4) Compliance-enhanced design: Based on explainable reports, decision transparency is improved by 89%.
[0065] The process shown in S101-S104 above uses the environmental benefit-privacy association matrix and the target reinforcement learning model to achieve a reasonable allocation of the privacy budget of green financial data, which is objective and efficient.
[0066] like Figure 2 As shown, it is a flow chart of a green financial data differential privacy budget allocation method based on reinforcement learning provided in an embodiment of the present application, which includes the following steps.
[0067] S201: Obtain local parameter information uploaded by each distributed node according to the federated training mechanism.
[0068] Among them, the federated training mechanism is used to use the target reinforcement learning model as the federated model, and the model parameters obtained by training the federated model with local data are used as local parameter information.
[0069] In some examples, the model parameters include horizontal federation parameters and vertical federation parameters, the horizontal federation parameters include at least model gradient parameters, and the vertical federation parameters include at least environmental benefit feature embedding vectors.
[0070] It should be noted that the federated training mechanism can be used to collaboratively train models on multiple distributed nodes while protecting data privacy.
[0071] S202: Utilizing each local parameter information, the parameters of the target reinforcement learning model are tuned to optimize the target reinforcement learning model.
[0072] Among them, using each local parameter information to tune the parameters of the target reinforcement learning model can not only optimize the parameters, but also ensure data privacy and security because each distributed node uploads local parameter information instead of data.
[0073] The process shown in S201-S202 above utilizes a federated learning architecture to optimize the performance of the target reinforcement learning model while ensuring data privacy and security.
[0074] like Figure 3 As shown, it is a flow chart of a green financial data differential privacy budget allocation method based on reinforcement learning provided in an embodiment of the present application, which includes the following steps.
[0075] S301: Based on the privacy budget of green financial data, generate corresponding privacy-preserving data, and publish the privacy-preserving data to multiple distributed nodes.
[0076] Among them, the distributed nodes pre-deploy a trusted execution environment.
[0077] S302: Based on the pre-deployed monitoring system, feedback data generated by each distributed node corresponding to the privacy-preserving data is obtained, and each feedback data is used as new training data to train the target reinforcement learning model to achieve an update of the target reinforcement learning model.
[0078] Among them, the monitoring system includes data layer, privacy layer, energy efficiency layer and compliance layer. Specifically, the data layer is used to track the fluctuations of green indicators in real time, the privacy layer is used to continuously evaluate the consumption rate of the privacy budget, the energy efficiency layer is used to monitor the consumption of resource costs, and the compliance layer is used to generate regulatory audit logs.
[0079] In some examples, the core purpose of deploying a monitoring system is to provide technical support for the adaptive allocation of privacy budgets. Generally speaking, the monitoring system is the "nerve center" for the implementation of differential privacy budgets. Its layered data perception and response capabilities provide an indispensable decision-making basis and control channel for the adaptive allocation of privacy budgets.
[0080] The process shown in S301-S302 above updates the target reinforcement learning model through multi-dimensional monitoring feedback from the monitoring system, thereby optimizing the privacy budget allocation reinforcement learning decision process.
[0081] like Figure 4 As shown, it is a flow chart of a green financial data differential privacy budget allocation method based on reinforcement learning provided in an embodiment of the present application, which includes the following steps.
[0082] S401: Determine the SHAP value of each feature during the learning process of the target reinforcement learning model.
[0083] Among them, the SHAP value is used to quantify the contribution of features to the privacy budget allocation results.
[0084] In some examples, a model interpretability method based on game theory (Shapley Additive Explanations, SHAP) can be used to obtain the SHAP value of each feature.
[0085] S402: Generate a corresponding interpretability report based on the SHAP value of each feature.
[0086] Among them, explainable reports are used to meet the transparency requirements of financial supervision.
[0087] The process shown in S401-S402 above uses the SHAP value of each feature to generate a corresponding explainability report to meet the transparency requirements of financial supervision.
[0088] like Figure 6 As shown, it is a schematic diagram of the architecture of a green financial data differential privacy budget allocation device based on reinforcement learning provided in an embodiment of the present application, including the units shown below.
[0089] The data collection unit 100 is used to perform format-preserving encryption on the green financial data collected in real time to obtain desensitized financial data.
[0090] The data processing unit 200 is used to input the desensitized financial data into a pre-trained generative adversarial network to obtain time series data of multiple features; the features are obtained by feature engineering based on historical green financial data.
[0091] The data association unit 300 is used to construct an environmental benefit-privacy association matrix based on the time series data of multiple features in combination with the association matrix method; the environmental benefit-privacy association matrix is used to quantify the contribution of multiple features to the privacy leakage risk.
[0092] The reinforcement learning unit 400 is used to input the contribution of multiple features and time series data into a pre-built target reinforcement learning model to obtain the privacy budget allocation result output by the target reinforcement learning model; the target reinforcement learning model uses a multi-objective reward function to dynamically adjust the weight coefficient to achieve Pareto optimality, and the target reinforcement learning model also uses an adaptive noise injection mechanism to achieve self-matching of noise intensity and data fluctuation period based on the reinforcement learning action output; the adaptive noise injection mechanism includes a Laplace noise dynamic scaler; the privacy budget allocation result includes the privacy budget of green financial data.
[0093] Optionally, the target reinforcement learning model includes a deep deterministic policy gradient model, which is used to model the privacy budget allocation as a continuous action space optimization problem.
[0094] Optionally, the multi-objective reward function includes a triple reward, which includes a privacy protection reward, a data utility reward, and a resource cost reward. The privacy protection reward is negatively correlated with the probability of privacy leakage, the data utility reward is positively correlated with the accuracy of time series data, and the resource cost reward is negatively correlated with the consumption of computing resources.
[0095] Optionally, a Laplacian noise dynamic scaler is trained based on a Long Short-Term Memory network.
[0096] The model update unit 500 is used to obtain the local parameter information uploaded by each distributed node according to the federated training mechanism; the federated training mechanism is used to use the target reinforcement learning model as the federated model, and use the model parameters obtained by training the federated model with local data as the local parameter information; the target reinforcement learning model is parameter-tuned using each local parameter information to achieve optimization of the target reinforcement learning model.
[0097] Optionally, the model update unit 500 is also used to generate corresponding privacy protection data based on the privacy budget of green financial data, and publish the privacy protection data to multiple distributed nodes; the distributed nodes pre-deploy a trusted execution environment; based on the pre-deployed monitoring system, obtain the feedback data generated by the privacy protection data at each distributed node, and use each feedback data as new training data to train the target reinforcement learning model to achieve the update of the target reinforcement learning model.
[0098] The report generation unit 600 is used to determine the SHAP value of each feature during the learning process of the target reinforcement learning model; the SHAP value is used to quantify the contribution of the feature to the privacy budget allocation result; based on the SHAP value of each feature, a corresponding explainability report is generated; the explainability report is used to meet the transparency requirements of financial supervision.
[0099] The above-mentioned units utilize the environmental benefit-privacy association matrix and the target reinforcement learning model to achieve a reasonable allocation of the privacy budget for green financial data, which is objective and efficient.
[0100] The present application also provides a computer-readable storage medium, which includes a stored program, wherein the program executes the green financial data differential privacy budget allocation method based on reinforcement learning provided by the present application.
[0101] This application also provides an electronic device comprising: a processor, a memory, and a bus. The processor and the memory are connected via the bus, the memory being used to store a program, and the processor being used to run the program. When the program runs, the method for allocating differentially private budgets for green financial data based on reinforcement learning provided in this application is executed.
[0102] Although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination.
[0103] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the disclosure herein is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned disclosure. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A green finance data differential privacy budget allocation method based on reinforcement learning, characterized by: include: Perform format-preserving encryption on green finance data collected in real time to obtain desensitized financial data; Inputting the desensitized financial data into a pre-trained generative adversarial network to obtain time series data of multiple features; The features are obtained through feature engineering based on historical green finance data; Based on the time series data of the multiple features, combined with the correlation matrix method, an environmental benefit-privacy correlation matrix is constructed; the environmental benefit-privacy correlation matrix is used to quantify the contribution of the multiple features to the privacy leakage risk; Based on the contribution of multiple features and time series data, they are input into a pre-built target reinforcement learning model to obtain the privacy budget allocation result output by the target reinforcement learning model; the target reinforcement learning model uses a multi-objective reward function to dynamically adjust the weight coefficient to achieve Pareto optimality, and the target reinforcement learning model also uses an adaptive noise injection mechanism to achieve self-matching of noise intensity and data fluctuation period based on the reinforcement learning action output; the adaptive noise injection mechanism includes a Laplace noise dynamic scaler; the privacy budget allocation result includes the privacy budget of the green financial data.
2. The method according to claim 1, characterized in that The method further comprises: Obtaining local parameter information uploaded by each distributed node according to a federated training mechanism; the federated training mechanism is used to use the target reinforcement learning model as a federated model and train the federated model using local data to obtain model parameters as the local parameter information; Utilize the local parameter information to perform parameter tuning on the target reinforcement learning model to achieve optimization of the target reinforcement learning model.
3. The method according to claim 1, characterized in that The method further comprises: Based on the privacy budget of the green financial data, corresponding privacy-preserving data is generated, and the privacy-preserving data is published to a plurality of distributed nodes; the distributed nodes are pre-deployed with a trusted execution environment; Based on a pre-deployed monitoring system, feedback data generated by each of the distributed nodes corresponding to the privacy-protected data is obtained, and each of the feedback data is used as new training data to train the target reinforcement learning model to achieve an update of the target reinforcement learning model.
4. The method according to claim 1, wherein The method further comprises: Determining a SHAP value for each of the features during the learning process of the target reinforcement learning model; the SHAP value is used to quantify the contribution of the feature to the privacy budget allocation result; Based on the SHAP value of each of the features, a corresponding explainability report is generated; the explainability report is used to meet the transparency requirements of financial supervision.
5. The method according to claim 1, wherein The target reinforcement learning model includes a deep deterministic policy gradient model, which is used to model privacy budget allocation as a continuous action space optimization problem.
6. The method according to claim 1, characterized in that The multi-objective reward function includes a triple reward, which includes a privacy protection reward, a data utility reward, and a resource cost reward. The privacy protection reward is negatively correlated with the probability of privacy leakage, the data utility reward is positively correlated with the accuracy of time series data, and the resource cost reward is negatively correlated with the consumption of computing resources.
7. The method according to claim 1, characterized in that The Laplace noise dynamic scaler is based on long short-term memory network training.
8. A green financial data differential privacy budget allocation device based on reinforcement learning, characterized in that: include: A data collection unit, used to encrypt the green financial data collected in real time in a format-preserving manner to obtain desensitized financial data; A data processing unit, configured to input the desensitized financial data into a pre-trained generative adversarial network to obtain time series data of multiple features; The features are obtained through feature engineering based on historical green finance data; A data association unit is configured to construct an environmental benefit-privacy association matrix based on the time series data of the plurality of features in combination with an association matrix method; the environmental benefit-privacy association matrix is configured to quantify the contribution of the plurality of features to the risk of privacy leakage; A reinforcement learning unit is used to input the contribution of multiple features and time series data into a pre-built target reinforcement learning model to obtain a privacy budget allocation result output by the target reinforcement learning model; the target reinforcement learning model adopts a multi-objective reward function to dynamically adjust the weight coefficient to achieve Pareto optimality, and the target reinforcement learning model also adopts an adaptive noise injection mechanism to achieve self-matching of noise intensity and data fluctuation period based on the reinforcement learning action output; the adaptive noise injection mechanism includes a Laplace noise dynamic scaler; the privacy budget allocation result includes the privacy budget of the green financial data.
9. A storage medium, characterized in that: The storage medium includes a stored program, wherein, when the program is run by a processor, the green financial data differential privacy budget allocation method based on reinforcement learning according to any one of claims 1 to 7 is executed.
10. An electronic device, characterized in that: include: processor, memory, and bus; The processor is connected to the memory via the bus; The memory is used to store programs, and the processor is used to run programs, wherein when the program is run by the processor, the green financial data differential privacy budget allocation method based on reinforcement learning according to any one of claims 1 to 7 is executed.