Task execution method and device for reinforcement learning model privacy protection, and medium
By constructing evaluation neural networks and target evaluation neural networks in deep reinforcement learning models, dynamically adjusting the privacy budget, and using the Laplace mechanism to add differential privacy, the balance of privacy protection and performance maintenance is solved, and effective protection of user privacy and stability of model performance is achieved.
Patent Information
- Application Number
- CN202510224992.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art, when protecting the privacy of deep reinforcement learning models, add noise causes a significant decline in model performance, making it difficult to find a balance between privacy protection and performance maintenance.
Using a MAPPO-based task execution method, by building an evaluation neural network and a target evaluation neural network, dynamically adjusting the privacy budget, using the Laplace mechanism to add differential privacy to the agent's state, realizing state encryption, and output task execution strategies.
Effectively protect user privacy, prevent sensitive information leakage, while maintaining stable model performance, not affecting data processing efficiency, and enhancing the security and robustness of the model.
Smart Images

Figure CN120277704A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of reinforcement learning, and in particular relates to a task execution method, device, and medium for privacy protection of reinforcement learning models. Background Art
[0002] With the rapid development of big data and artificial intelligence technology, deep reinforcement learning has been widely used in many fields, such as autonomous driving, drone control, etc. However, reinforcement learning models need to process a large amount of user data during training to continuously improve their decision-making and learning capabilities. However, this data often contains sensitive information of users, which inevitably leads to the risk of user privacy leakage.
[0003] Differential privacy technology, as an emerging privacy protection method, provides a possible solution to this problem. It adds noise during data processing, making it difficult for attackers to infer information about the original data from the processed data. In this way, even if the data is leaked, it is difficult for attackers to easily spy on the user's privacy.
[0004] However, traditional differential privacy methods face a huge challenge when dealing with deep reinforcement learning models. Due to the addition of additional noise, the performance of the model is often significantly affected when processing data. Therefore, how to maintain the performance of the reinforcement learning model while protecting privacy has become a difficult problem that needs to be solved urgently. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides a task execution method, device, and medium for privacy protection of reinforcement learning models.
[0006] In a first aspect, an embodiment of the present invention provides a task execution method for privacy protection of a reinforcement learning model, the method comprising:
[0007] Build mission execution scenarios based on MAPPO;
[0008] Receive the current state of the agent decision model in the task execution scenario, calculate the similarity between the strategy made by the agent decision model and the recommended strategy according to the current state, and thus update the state; construct an evaluation neural network, wherein the evaluation neural network is used to estimate the privacy budget according to the updated state; construct a target evaluation neural network, wherein the target evaluation neural network has the same network architecture as the evaluation neural network, and the target evaluation neural network and the evaluation neural network are collaboratively updated and trained to describe the maximum privacy budget;
[0009] The target evaluation neural network outputs the maximum privacy budget based on the current state, and uses the Laplace mechanism to add differential privacy to the state of the agent according to the maximum privacy budget to achieve state encryption;
[0010] The agent outputs a task execution policy through state encryption to execute the task.
[0011] In a second aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, the memory being coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned task execution method for privacy protection of the reinforcement learning model.
[0012] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned task execution method for privacy protection of the reinforcement learning model is implemented.
[0013] In a fourth aspect, an embodiment of the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the above-mentioned task execution method for privacy protection of the reinforcement learning model is implemented.
[0014] Compared with the prior art, the beneficial effects of the present invention are:
[0015] Based on the differential privacy deep reinforcement learning model, the present invention adds noise to the training data, making it difficult for attackers to infer users' sensitive information by analyzing the model output or training data. This addition of noise is carried out on the premise of ensuring the model performance. This method not only effectively protects user privacy but also does not significantly affect the model performance.
[0016] The differential privacy method of the present invention does not require data decryption for calculation, which greatly improves the data processing efficiency. At the same time, because the differential privacy technology adds carefully designed noise to the data, even if the attacker has specific data about a particular individual, it is almost impossible to infer specific information about the individual. This technology can greatly enhance the privacy protection ability of the deep reinforcement learning model and effectively prevent the leakage and abuse of user data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0018] Figure 1 It is a flowchart of the task execution method for privacy protection of the reinforcement learning model provided by the embodiment of the present invention;
[0019] Figure 2 Schematic diagram of the scenario of multi-agent game confrontation of drones provided by the embodiments of the present invention;
[0020] Figure 3 Flowchart of the task execution method for privacy protection of the reinforcement learning model in the drone game confrontation scenario provided by the embodiments of the present invention;
[0021] Figure 4 Schematic diagram of privacy protection based on differential privacy in the drone game confrontation scenario provided by the embodiments of the present invention;
[0022] Figure 5 Schematic diagram of an electronic device provided by the embodiments of the present invention. Specific embodiments
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] It should be noted that, without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0025] As Figure 1 shown, the embodiments of the present invention provide a task execution method for privacy protection of the reinforcement learning model. The method includes the following steps:
[0026] Step S1, constructing a task execution scenario based on MAPPO, including: setting the number of agents, the joint system state of multiple agents, the actions of agents, the states of agents, the state transition function, and the reward function.
[0027] Step S2, receiving the current state of the agent decision model in the task execution scenario, calculating the similarity between the policy made by the agent decision model and the recommended policy according to the current state, so as to update the state; constructing an evaluation neural network, where the evaluation neural network is used to estimate the privacy budget according to the updated state; constructing a target evaluation neural network, where the target evaluation neural network has the same network architecture as the evaluation neural network, and the target evaluation neural network and the evaluation neural network are jointly updated and trained to obtain the maximum privacy budget.
[0028] The expression for calculating the similarity between the policy made by the agent decision model and the recommended policy is as follows:
[0029]
[0030] Wherein, represents the similarity, v p,j represents the strategy made by the agent decision model, v i,j represents the recommended strategy, i represents the i-th agent, j represents the j-th strategy dimension, B represents the strategy dimension, and H represents the number of agents.
[0031] It should be noted that in this example, by observing the strategy behavior, data flow, and privacy leakage of the agent, when a high privacy leakage risk is found, the scheme will automatically and appropriately reduce the privacy budget to strengthen the protection; on the contrary, when the privacy protection is sufficient but the learning performance is limited, the privacy budget will be appropriately increased to optimize the learning effect.
[0032] The process of adding differential privacy to the state of the agent by using the Laplace mechanism according to the maximum privacy budget includes:
[0033] The distribution expression of Laplace is as follows:
[0034]
[0035] The expression of sensitivity is as follows:
[0036] S(f) = max||f(D) - f(D')||1
[0037] Wherein, f(D) represents the decision before adding the privacy budget, and f(D') represents the decision after adding the privacy budget;
[0038]
[0039] Wherein, ε is the privacy budget.
[0040] Step S3, the target evaluation neural network outputs the maximum privacy budget based on the current state, and adds differential privacy to the state of the agent by using the Laplace mechanism according to the maximum privacy budget to achieve state encryption.
[0041] Step S4, the agent with state encryption outputs a task execution strategy to execute the task.
[0042] Furthermore, the method further includes performance evaluation of the trained target evaluation neural network, and the expression is as follows:
[0043]
[0044] Wherein, u (k) represents the utility degree of this encryption method, represents the similarity of the r-th agent strategy, c1 represents the budget coefficient, x (k) represents the expected privacy budget, Denote the sensitivity level of the p dimension recommended in time period k, c2 represents the sensitivity coefficient, and ζ (k) represents the privacy metric.
[0045] It should be noted that the present invention can dynamically adjust the privacy budget according to the actual situation during the learning process. This process can be achieved by observing the behavior of the agent, data flow, and privacy leakage. When the risk of privacy leakage is found to be high, the scheme will automatically and appropriately reduce the privacy budget to enhance protection; conversely, when the privacy protection is sufficient but the learning performance is limited, the privacy budget will be appropriately increased to optimize the learning effect.
[0046] Embodiment 1
[0047] Next, taking the UAV game confrontation scenario as an example, this example illustrates the specific implementation process of the task execution method for privacy protection of the reinforcement learning model provided by the present invention.
[0048] In this example, in the UAV game confrontation scenario, deep reinforcement learning is used to generate a privacy budget through the expert data of the target UAV model, and then the target UAV is encrypted through differential privacy to achieve privacy protection. Figure 2 is the UAV game scenario used in this example. By training the optimal strategy on our side, accurate and effective strikes on enemy UAVs can be achieved. As Figure 3 shown, it specifically includes the following steps:
[0049] Step S1, build a UAV game confrontation scenario based on MAPPO, including: setting the number of agents, the joint system state of multiple agents, the action set of agents, the state space of agents, the state transition function, and the reward function.
[0050] Specifically, the step S1 includes:
[0051] This UAV game confrontation scenario essentially involves real-time agent decision-making problems in continuous action and state spaces. In this example, the variables used in the agent decision-making problem are summarized in a tuple, and the expression is as follows:
[0052] (N, S, A1, A2,..., A N , F, γ, R1, R2,..., R N )
[0053] In the formula, N is the number of agents; S is the joint system state of multiple agents; A1, A2,..., A N is the action set of agents; F is the state transition function, that is, according to the current state and joint action, give the probability distribution of the next state, F: S × A1 ×... × A n × S, → [0, 1]; R i(S, A1,..., A N , S) represents the reward for the agent to execute a joint action and reach the next state when in state S; γ represents the discount factor, where γ ∈ [0, 1].
[0054] Among them, the state space of the agent is as follows:
[0055]
[0056] In the formula, o red represents the state of the red - side UAV, o blue represents the state of the blue - side UAV, represents the distance between the red - side UAV and its friendly aircraft, represents the angle between the red - side UAV and its friendly aircraft, represents the distance between the red - side UAV and the enemy aircraft, represents the angle between the red - side UAV and the enemy aircraft, represents the flight altitude of the red - side UAV, represents the survival state of the red - side UAV. i represents the i - th UAV.
[0057] Among them, the action space of the agent is set as follows:
[0058] The movement of the UAV depends on the precise adjustment of speed and heading angle, and firing is the key means to strike the enemy target. When making a decision, each agent faces two choices: firing and moving. Once firing is selected, the movement direction and speed of the UAV will maintain the previous state to ensure the coherence and accuracy of the strike; when moving is selected, the new movement direction and speed need to be clearly specified to adapt to the changes in the UAV game confrontation situation. When the enemy UAV enters the effective attack range of our agent, our agent will, based on the battlefield situation and its own strategy, choose to fire with a probability of p fire . After firing, our UAV also has a probability of p br of successfully shooting down the enemy aircraft.
[0059] Among them, the expression of the reward function is as follows:
[0060]
[0061] It should be noted that in the process of training our agent, the reinforcement learning method is adopted, while the enemy uses a fixed training method. In order to achieve the optimization of decision - making, the goal of this example is to ensure that the number of our agents is as large as possible and to eliminate as many enemy agents as possible. Therefore, when our side successfully destroys an enemy agent, the reward value is set to a positive value; on the contrary, when our agent is destroyed by the enemy, the reward value is negative. When all the agents of any side are destroyed, the global minimum value and the global maximum value will be given as rewards and punishments in extreme cases.
[0062] Considering the cooperation among multiple agents in this scenario and the complexity of communication issues, this example specifically designs a distance-based sparse reward mechanism. This mechanism aims to encourage the agents within the cluster to maintain an appropriate distance, thereby optimizing the overall tactical layout and action coordination. At the same time, for agents that exceed the restricted flight area, we will also impose corresponding penalties to ensure the orderly and efficient actions of the entire team.
[0063] By setting such a reward function, it can not only guide the agents to make more reasonable decisions, but also promote their close cooperation and effective communication, thereby improving the combat effectiveness of the entire team.
[0064] Step S2: Receive the current state of the agent decision model in the task execution scenario, calculate the similarity between the policy made by the agent decision model and the recommended policy based on the current state, so as to update the state; construct an evaluation neural network, which is used to estimate the privacy budget according to the updated state; construct a target evaluation neural network, the network architecture of the target evaluation neural network is the same as that of the evaluation neural network, and the target evaluation neural network and the evaluation neural network are jointly updated and trained to obtain the maximum privacy budget.
[0065] Specifically, the state space of the agent UAV decision model includes the privacy requirements of the current UAV state The similarity between two UAV states, the current state of the UAV [v i,j 1≤i≤H,1≤j≤B , and the privacy metric ζ. Define the state space as composed of all feasible privacy requirements, similarities, and privacy metrics of S, and the expression is as follows:
[0066] S = {[z p , [μ i 1≤i≤H , ζ]: z p ∈ {0, 1,..., Z}, 0 ≤ μ i ≤ 1, ζ ∈ {0, 1}}
[0067] In the formula, z p represents the sensitivity level of the recommended policy p, Z represents the maximum sensitivity level, μ i represents the cosine similarity of the policy, and H represents the recommended policy.
[0068] The action space of the agent UAV decision model includes: after inputting the privacy-protected data or the original data without privacy protection into the agent UAV decision model, the decisions and their distributions given by the agent UAV decision model, including instructions such as acceleration and deceleration, changing the nose direction, locking the target, or launching an attack, and their specific parameters.
[0069] It should be noted that the design of the action space should ensure that it can ultimately determine whether the privacy protection process has significantly interfered with the UAV decision-making model, so as to find the best balance between privacy protection and task completion. Ideally, the data after privacy protection should not cause visible changes in the action space compared to the original data.
[0070] The reward function of the intelligent agent UAV decision-making model follows the following design principles: when the privacy budget adjustment strategy can effectively protect the UAV state and generate high-quality decisions, a higher first reward is given; when the privacy budget adjustment strategy leads to the leakage of the UAV privacy state or the decline of the UAV model decision-making quality, a lower second reward is given. In this way, the reward function can guide the reinforcement learning model to find a balance between protecting user privacy and providing high-quality recommendations, and then provide good privacy protection for the UAV state.
[0071] It should be noted that the reward function is the standard used to evaluate the quality of actions in reinforcement learning. In this example, two factors, namely privacy protection and decision interference degree, need to be comprehensively considered.
[0072] The expression for calculating the similarity between the strategy made by the intelligent agent decision-making model and the recommended strategy is as follows:
[0073]
[0074] In the formula, represents the similarity, v p,j represents the strategy made by the intelligent agent decision-making model, v i,j represents the recommended strategy, i represents the i-th intelligent agent, j represents the j-th strategy dimension, B represents the strategy dimension, and H represents the number of intelligent agents.
[0075] It should be noted that in this example, by observing the strategy behaviors, data flows, and privacy leakage situations of the intelligent agents, when the risk of privacy leakage is found to be relatively high, the scheme will automatically and appropriately reduce the privacy budget to strengthen protection; conversely, when the privacy protection is sufficient but the learning performance is limited, the privacy budget will be appropriately increased to optimize the learning effect.
[0076] The process of adding differential privacy to the state of the intelligent agent according to the maximum privacy budget using the Laplace mechanism includes:
[0077] The distribution expression of Laplace is as follows:
[0078]
[0079] The expression for the sensitivity is as follows:
[0080] S(f) = max||f(D) - f(D')||1
[0081] In the formula, f(D) represents the decision before adding the privacy budget, and f(D') represents the decision after adding the privacy budget;
[0082]
[0083] In the formula, ε is the privacy budget.
[0084] Step S3, the target evaluation neural network outputs the maximum privacy budget based on the current state, and uses the Laplace mechanism to add differential privacy to the state of the intelligent agent according to the maximum privacy budget to achieve state encryption.
[0085] Step S4, the intelligent agent with state encryption outputs a task execution strategy, thereby executing the task.
[0086] Furthermore, the method further includes performance evaluation of the trained target evaluation neural network, and the expression is as follows:
[0087]
[0088] In the formula, u (k) represents the utility degree of this encryption method, represents the similarity of the r-th intelligent agent strategy, c1 represents the budget coefficient, and x (k) represents the expected privacy budget, represents the sensitivity level of recommending the p dimension in period k, c2 represents the sensitivity coefficient, and ζ (k) represents the privacy metric.
[0089] Furthermore, this example formulates a privacy-aware recommendation game to evaluate the upper bound of the performance of the DUPP scheme provided by this example, that is, the convergence performance of DUPP. The user equipment selects the privacy budget x (k) ∈(0,X) to optimize the utility u (k) , and at the same time selects the attack probability y (k) to reduce the utility of the user equipment. Assume that the items recommended by H have the same sensitivity level Z.
[0090] The privacy loss of the user equipment is defined as Then correspondingly, the privacy protection level of the user equipment is The recommendation quality ρ (k) is approximately defined according to the similarity between the characteristics of the decision r generated after adding privacy protection and the characteristics of the decision p generated by the original state:
[0091]
[0092] During the training of the constructed intelligent agent UAV decision-making model, we need to flexibly adjust the privacy budget ε and other relevant parameters according to the actual situation. This step aims to find the best balance between privacy protection and model performance. To achieve this goal, we have conducted multiple experiments and validations, continuously iterating and optimizing the parameter settings of the model to reach the optimal solution that not only meets the privacy protection requirements but also ensures the model performance.
[0093] In a UAV game confrontation scenario as Figure 2 shown, this example is committed to achieving precise and effective strikes on the blue side's UAV and other weapon systems. To achieve this goal, we adopt the multi-agent proximal policy optimization method to strictly train the red side's UAV. During the game process, the red side's UAV shows excellent mobility and precision by continuously adjusting its flight altitude, roll attitude, and pitch attitude, successfully shooting down the enemy UAV and achieving the minimum loss in the process.
[0094] In summary, the present invention provides a task execution method for privacy protection of a reinforcement learning model, which aims to ensure that the data privacy of the model is not leaked in a complex game environment while maintaining the stability of its performance. By introducing differential privacy technology, we inject random noise during the training process, making it difficult for attackers to infer the original training data by observing the output of the model, thereby protecting the security of sensitive information. In addition, we combine deep reinforcement learning algorithms so that the model can still learn effective strategies while protecting privacy. This method not only improves the security of the model but also enhances its robustness in practical applications. Through continuous testing and optimization, we can discover and fix potential privacy leakage risks to ensure the stable operation of the multi-agent game reinforcement learning model under the premise of ensuring privacy.
[0095] Correspondingly, the present application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the task execution method for privacy protection of a reinforcement learning model as described above. As Figure 5 shown, this is a hardware structure diagram of a device with any data processing ability where the task execution method for privacy protection of a reinforcement learning model provided by an embodiment of the present invention is located. In addition to Figure 5 the processors, memory, and network interfaces shown, any device with data processing ability where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with any data processing ability, which will not be elaborated here.
[0096] Correspondingly, the present application also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the task execution method for privacy protection of the reinforcement learning model as described above is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0097] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary.
[0098] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A task execution method for privacy protection of reinforcement learning models, characterized in that, The method includes: Constructing a task execution scenario based on MAPPO; Receiving the current state of the agent decision model in the task execution scenario, calculating the similarity between the policy made by the agent decision model and the recommended policy according to the current state, so as to update the state; constructing an evaluation neural network, where the evaluation neural network is used to estimate the privacy budget according to the updated state; constructing a target evaluation neural network, where the target evaluation neural network has the same network architecture as the evaluation neural network, and the target evaluation neural network and the evaluation neural network are jointly updated and trained to obtain the maximum privacy budget; The target evaluation neural network outputs the maximum privacy budget based on the current state, and adds differential privacy to the state of the agent according to the maximum privacy budget by using the Laplace mechanism to achieve state encryption; The agent with encrypted state outputs a task execution policy, thereby executing the task.
2. The task execution method for privacy protection of a reinforcement learning model according to claim 1, wherein Constructing a task execution scenario based on MAPPO includes: setting the number of agents, the joint system state of multiple agents, the actions of the agents, the states of the agents, the state transition function, and the reward function.
3. The task execution method for privacy protection of a reinforcement learning model according to claim 1, wherein, The expression for calculating the similarity between the policy made by the agent decision model and the recommended policy is as follows: In the formula, represents the similarity, v p,j represents the strategy made by the agent decision model, v i,j represents the recommended strategy, i represents the i-th agent, j represents the j-th strategy dimension, B represents the strategy dimension, and H represents the number of agents.
4. A task execution method for privacy protection of a reinforcement learning model according to claim 1, characterized in that, The process of adding differential privacy to the state of the agent according to the maximum privacy budget by using the Laplace mechanism includes: The distribution expression of Laplace is as follows: The expression for sensitivity is as follows: S(f) = max||f(D) - f(D')||1 In the formula, f(D) represents the decision before adding the privacy budget, and f(D') represents the decision after adding the privacy budget; In the formula, ε is the privacy budget.
5. A task execution method for privacy protection of a reinforcement learning model according to claim 1, characterized in that, The method further includes performing performance evaluation on the trained target evaluation neural network, and the expression is as follows: where u (k) represents the utility degree of this encryption method, represents the similarity of the r-th agent strategy, c1 represents the budget coefficient, x (k) represents the expected privacy budget, represents the sensitivity level of recommending the p dimension in period k, c2 represents the sensitivity coefficient, ζ (k) represents the privacy metric.
6. An electronic device, comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the task execution method for privacy protection of the reinforcement learning model according to any one of claims 1-5 above.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the task execution method for privacy protection of the reinforcement learning model according to any one of claims 1-5.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, it implements the task execution method for privacy protection of the reinforcement learning model according to any one of claims 1-5.