Encryption mechanism dynamic optimization method and system based on limited Markov decision process

Through the dynamic optimization method of encryption mechanism based on the restricted Markov decision-making process, the existing encryption methods are solved in the shortage of complex key management and high computing resource consumption, and adaptive optimization of encryption mechanism and dynamic adjustment of security policies are realized, which significantly improves security and efficiency.

CN120012127APending Publication Date: 2025-05-16STATE GRID JIANGSU ELECTRIC POWER CO LTD SUZHOU BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510020795.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing encryption methods are difficult to effectively respond to new threats and challenges when facing the complexity of key management, high computing resources consumption and changing security environments.

Method used

The encryption mechanism dynamic optimization method based on the restricted Markov decision-making process is adopted. By constructing the restricted Markov decision-making process, building an anti-autocoder to evaluate the security of the encryption mechanism, and combining the risk cost function to divide the encryption mechanism into a secure set and a violation set. Finally, the encryption mechanism that is not feasible and centralized is optimized based on the feasibility of the encryption mechanism, and the security policy is dynamically adjusted.

Benefits of technology

Adaptive optimization of the encryption mechanism is achieved, rewards are maximized and costs do not exceed thresholds, significantly improving security and efficiency, and adapting to a changing security environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012127A_ABST
    Figure CN120012127A_ABST
Patent Text Reader

Abstract

The invention relates to an encryption mechanism dynamic optimization method and system based on a limited Markov decision process, and the method comprises the steps: constructing the limited Markov decision process based on a to-be-optimized encryption mechanism, and the limited Markov decision process comprises a risk cost function; constructing an adversarial automatic encoder to evaluate the security of the encryption mechanism, and dividing the encryption mechanism into a security set and a violation set in combination with a risk cost function; constructing an optimal feasibility function to divide the security set and the violation set into a feasible set and a non-feasible set; the encryption mechanism in the infeasible set is optimized based on the security policy of the encryption mechanism feasibility, and the security policy is dynamically adjusted. According to the method, the overall safety and stability can be always kept in the optimization process of the encryption mechanism; by designing an optimization strategy adjustment method, the security and performance of the encryption mechanism are gradually improved in a continuous strategy optimization process, and finally self-learning and dynamic adjustment of an encryption mechanism optimization strategy are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power communication data network security, and more specifically, to a method and system for dynamically optimizing an encryption mechanism based on a restricted Markov decision process. Background Art

[0002] In the digital age, information has become an indispensable key asset in all fields. The security of information in the field of power communication is particularly important. Once sensitive data is leaked or used maliciously, it may cause serious property losses and bring security risks. Therefore, ensuring the privacy and security of data during transmission and storage is of decisive significance for building a secure digital environment. However, with the continuous expansion of the scale of power data networks, the continuous increase in business, the continuous changes in the environment and the continuous evolution of attack methods, traditional static security measures have been unable to cope with new threats and challenges.

[0003] Traditional encryption methods are mainly divided into two categories: symmetric encryption and asymmetric encryption, each of which is adaptive to different security requirements. Symmetric encryption methods, such as AES (Advanced Encryption Standard) and Chacha20, dominate large-scale data processing with their excellent processing speed and efficiency. These methods use the same key for encryption and decryption operations, which has the advantage of being able to quickly process large amounts of data, but also face the challenges of key distribution and management. Asymmetric encryption methods, such as RSA and elliptic curve cryptography (ECC), use a pair of public and private keys to operate, support digital signatures and secure key exchange, and are very suitable for use in open network environments. In addition, homomorphic encryption technology, as an innovative method, allows data processing directly in the ciphertext state, which provides new possibilities for cloud computing and data outsourcing services. At the same time, quantum key distribution (QKD) and quantum computing-oriented anti-quantum encryption methods (Post-quantum Cryptography) represent the latest developments in encryption technology, which are designed to cope with the security challenges that may be brought about by future quantum computers. These methods each have their own advantages and limitations, and their design and implementation complexity pose new challenges for the further development of encryption technology.

[0004] Although existing encryption methods provide a solid guarantee for data security, they each have a series of limitations. Symmetric encryption methods such as AES are excellent in processing speed, but the complexity and security of key management remain a major challenge, especially in large-scale systems, where the distribution and storage of keys require additional security measures. Asymmetric encryption methods, such as RSA and ECC, provide a more secure way of exchanging data in a public environment, but their processing efficiency and computing resource consumption are relatively high, which limits their application on resource-constrained devices. In addition, with the progress of quantum computing, traditional asymmetric encryption methods face the risk of being cracked by quantum attacks. Although homomorphic encryption technology provides the ability to perform calculations in a ciphertext state, the current implementation method still needs to be improved in terms of efficiency and practicality. Although quantum encryption technology provides extremely high security in theory, there are problems such as strong dependence on equipment and environment and high implementation cost in practical applications, which limits its widespread deployment. Summary of the invention

[0005] In order to address the deficiencies in the prior art, the present invention provides a dynamic optimization method for an encryption mechanism based on a restricted Markov decision process, which can effectively address the technical problems of key management complexity and high consumption of computing resources by dynamically evaluating and adjusting encryption strategies, and adapt to the ever-changing security environment.

[0006] The present invention adopts the following technical solution.

[0007] A method for dynamic optimization of encryption mechanism based on restricted Markov decision process comprises the following steps:

[0008] Based on the encryption mechanism to be optimized, a restricted Markov decision process is constructed, and the restricted Markov decision process includes a risk cost function;

[0009] Construct an adversarial autoencoder to evaluate the security of the encryption mechanism, and divide the encryption mechanism into a safe set and a violation set in combination with the risk cost function;

[0010] Construct the optimal feasibility function to divide the safety set and violation set into feasible set and infeasible set;

[0011] The security strategy based on the feasibility of the encryption mechanism optimizes the encryption mechanism of the infeasible concentration and dynamically adjusts the security strategy.

[0012] Preferably, the restricted Markov decision process includes a state space, an action space, a state transition function, a reward function, a risk cost function and a discount factor;

[0013] Among them, the state space S represents the current encryption mechanism state, covering various parameters of the encryption mechanism state and security situation indicators;

[0014] The action space A represents the possible policy adjustment operations that the encryption mechanism may take in each state;

[0015] The state transition function P(s′|s, a) describes the probability of transitioning from the current encryption mechanism state s∈S to the next encryption mechanism state s′∈S after taking action a∈A;

[0016] The reward function r(s, a) is used to evaluate the impact of each encryption mechanism state transition on security and performance;

[0017] The security loss function h is used to map the encryption mechanism state s∈S to a non-negative real value, which is used to evaluate the potential risk cost in the encryption mechanism state.

[0018] Preferably, the constructing of the adversarial autoencoder specifically comprises:

[0019] The adversarial autoencoder includes an autoencoder and an adversarial network, wherein the autoencoder includes an encoder and a decoder, the encoder is used to map the input encryption mechanism state into a low-dimensional latent space, and the decoder is used to obtain the reconstructed encryption mechanism state from the low-dimensional latent space; the adversarial network includes a discriminator, which is used to distinguish the original encryption mechanism state from the encryption mechanism state reconstructed by the autoencoder.

[0020] Preferably, the risk cost function specifically includes:

[0021] The risk cost function is used to calculate the risk cost h(s) of the encryption mechanism s:

[0022]

[0023] Where D(s) is the score output of the discriminator for the encryption mechanism state s, is the discriminator to reconstruct the encryption mechanism , ε is the risk tolerance factor, max{} is the maximum value operation, and w(s) represents the importance of the encryption mechanism state s;

[0024] Preferably, the encryption mechanism is divided into a safe set and a violation set in combination with a risk cost function, specifically including:

[0025] According to the risk cost function h(s), the encryption mechanism s is judged to belong to the security set or the violation set:

[0026] S s : ={s∈S:h(s)=0}

[0027] S v : ={s∈S:h(s)>0}

[0028] Where S represents the state space of the restricted Markov decision process, Ss represents the safe set, S v Represents the offending set.

[0029] Preferably, by measuring the impact of the encryption mechanism state on the reward, the importance weight w(s) of the encryption mechanism state s is calculated as follows:

[0030]

[0031] Among them, a1 and a2 represent two actions in the action space A of the restricted Markov decision process, and r(s, a1) and r(s, a2) represent the reward functions of the encryption mechanism state s under actions a1 and a2 respectively.

[0032] Preferably, the constructing of the optimal feasibility function H*(s) specifically includes:

[0033]

[0034] in, is the instantaneous violation indicator function, when the encryption mechanism state s at time t t When it is not in the security set, its value is 1; when the encryption mechanism state s at time t t Its value is 0 when in the security concentration; τ~π*, P(s) means starting from the encryption mechanism state s, using the optimal security strategy π* to sample the trajectory τ under the state transition function P, For violation set.

[0035] Preferably, dividing the safety set and the violation set into a feasible set and an infeasible set specifically includes:

[0036] According to the calculated value of the optimal feasibility function H*(s), the safety set and the violation set are divided into feasible sets and infeasible sets:

[0037]

[0038] in, represents the feasible set, Represents an infeasible set.

[0039] Preferably, the security strategy based on the feasibility of the encryption mechanism optimizes the encryption mechanism of the infeasible centralization, specifically including:

[0040] Combined with the actor-critic algorithm, the encryption mechanism in the infeasible set is guided to return to the feasible set by minimizing the maximum possible violation cost in the future, and the encryption mechanism in the feasible set is optimized by maximizing the expected reward. The resulting security strategy process is as follows:

[0041]

[0042] Where π represents the strategy, a~π(·|s) represents the sampling action a according to the strategy, Q(s, a) is the state-action value function, and Q c (s, a) is the state-action cost value function, H*(s) is the optimal feasibility function, λ is the Lagrange multiplier, d0 is the initial encryption mechanism state distribution, s~d0 represents the encryption mechanism state sampled from the initial encryption mechanism state distribution, and c is the threshold constant.

[0043] Preferably, the security policy is dynamically adjusted, specifically including:

[0044] The network parameters ω of the state-action value function Q(s, a), the state-action cost value function Q c The network parameter p of (s, a), the parameter θ of the policy network π, and the Lagrange multiplier λ are dynamically updated as follows:

[0045]

[0046]

[0047] Among them, s t 、s t+1 Respectively represent the encryption mechanism status at time t and time t+1, a t 、a t+1 Respectively represent the actions on the encryption mechanism state at time t and time t+1, represents the gradient, They represent the gradient of parameters ω, p, and λ respectively, ζ1, ζ2, ζ3, and ζ4 are the learning rates of the network respectively, and π θ represents a policy with parameter θ.

[0048] The present invention also proposes a dynamic optimization system for encryption mechanism based on restricted Markov decision process, which is used to implement the dynamic optimization method for encryption mechanism based on restricted Markov decision process, including: a restricted Markov decision process building module, a primary partitioning module, a secondary partitioning module and an optimization adjustment module;

[0049] The restricted Markov decision process building module is used to build a restricted Markov decision process based on the encryption mechanism to be optimized, and the restricted Markov decision process includes a risk cost function;

[0050] The initial partitioning module evaluates the security of the encryption mechanism by constructing an adversarial autoencoder and divides the encryption mechanism into a secure set and a violation set in combination with a risk cost function;

[0051] The secondary partitioning module divides the safety set and violation set into feasible set and infeasible set by constructing the optimal feasibility function;

[0052] The optimization and adjustment module optimizes the encryption mechanism of the infeasible concentration based on the security policy of the feasibility of the encryption mechanism and dynamically adjusts the security policy.

[0053] The present invention also provides a terminal, comprising a processor and a storage medium;

[0054] The storage medium is used to store instructions;

[0055] The processor is used to operate according to the instructions to execute the steps of the encryption mechanism dynamic optimization method based on the restricted Markov decision process.

[0056] The present invention also proposes a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the encryption mechanism dynamic optimization method based on a restricted Markov decision process are implemented.

[0057] The beneficial effect of the present invention is that, compared with the prior art, the present invention reconstructs the encryption mechanism by constructing an adversarial autoencoder to evaluate the confusability of the tampered encryption mechanism with the original mechanism, thereby quantifying the security of the encryption mechanism in the face of tampering. In order to ensure the security of the entire encryption mechanism adjustment process, the present invention introduces an optimal feasibility function to evaluate the feasibility of the encryption mechanism adjustment sequence from the initial state to the terminal state, ensuring that the overall security and stability are always maintained during the optimization process of the encryption mechanism. At the same time, the present invention designs a corresponding optimization strategy adjustment method for encryption mechanisms that are in the feasible set and those that are not in the feasible set, so that the encryption mechanism gradually improves its security and performance in the continuous strategy optimization process, and finally realizes the self-learning and dynamic adjustment of the encryption mechanism optimization strategy.

[0058] (1) Based on the restricted Markov decision process framework, the present invention defines the state space and action space of the encryption mechanism, optimizes the encryption mechanism by dividing the encryption mechanism into feasible set and infeasible set using the risk cost function through reward and risk cost function, and realizes the adaptive optimization of the encryption mechanism, maximizing the reward and ensuring that the cost does not exceed the threshold. The method can dynamically select and adjust the encryption strategy according to the changes in the security environment, significantly improving security and efficiency.

[0059] (2) The present invention constructs a security assessment of the encryption mechanism against the automatic encoder. By reconstructing the encryption mechanism and detecting its tampering resistance, the security of the encryption mechanism in the dynamic adjustment process is quantified. The encryption mechanism state is treated differently by the state importance weight. This method can effectively identify high-risk states and provide accurate decision support for the optimization of the encryption mechanism.

[0060] (3) The present invention evaluates the security of the encryption mechanism optimization path through the optimal feasibility function to ensure the feasibility and security of the entire optimization process. According to the feasibility results, different optimization strategies for feasible and infeasible set encryption mechanisms are designed, and the optimization strategies are continuously updated and adjusted, which can maximize the efficiency while ensuring security and quickly guide the unsafe state back to the safe range. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a flow chart of a method for dynamic optimization of an encryption mechanism based on a restricted Markov decision process in the present invention;

[0062] Figure 2 It is a detailed framework diagram of the dynamic optimization method of the encryption mechanism based on the restricted Markov decision process in the present invention;

[0063] Figure 3 is a network structure diagram of the actor-critic method in the present invention;

[0064] Figure 4 It is a structural diagram of the dynamic optimization system of the encryption mechanism based on the restricted Markov decision process in the present invention. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The embodiments described in this application are only part of the embodiments of the present invention, not all of them. Based on the spirit of the present invention, other embodiments obtained by ordinary technicians in this field without creative work are all within the scope of protection of the present invention.

[0066] like Figure 1 , 2 As shown, the present invention proposes a dynamic optimization method for an encryption mechanism based on a restricted Markov decision process, the method comprising the following steps:

[0067] Step 1, based on the encryption mechanism to be optimized, construct a restricted Markov decision process, where the restricted Markov decision process includes a risk cost function;

[0068] Specifically, the restricted Markov decision process constructed in the present invention is defined as μ:=(S, A, P, r, h, γ);

[0069] The restricted Markov decision process includes state space S, action space A, state transfer function P, reward function r, risk cost function h and discount factor γ;

[0070] The restricted Markov decision process constructed by the present invention is used to select and optimize a variety of known encryption mechanisms (such as symmetric encryption, asymmetric encryption, hash functions, etc.) in the initial state space. By exploring the parameters and strategies of the encryption mechanism, a new selectable encryption mechanism can be generated. Through the exploration and evaluation of the restricted Markov decision process, the present invention starts from these known encryption mechanisms, adjusts the encryption parameters and strategies (including key length, method type, encryption mode, etc.), continuously generates new encryption mechanisms, and expands the state space. This process continues until the termination state is reached (such as reaching the maximum number of exploration steps or finding the optimal encryption mechanism state), realizing the optimization and innovation of the encryption mechanism to improve the overall security and performance.

[0071] State space S: State space S represents the current state of the encryption mechanism, including various encryption mechanism parameters (such as encryption method type, key length, character set complexity) and security situation indicators (such as recent attack frequency, attack type, system vulnerability analysis results, etc.). Each state represents a specific encryption configuration and its security environment characteristics.

[0072] Action space A: Action space A includes the possible policy adjustment actions for the encryption mechanism in each state. For example, changing the encryption method type, adjusting the key length, modifying encryption parameters, adding multi-factor authentication, or adjusting the update frequency of the encryption policy. Each action selection will result in an update of the encryption mechanism and change the encryption mechanism state.

[0073] State transfer function P: The state transfer function P(s′|s, a) describes the probability of transitioning from the current state s∈S to the next state s′∈S after taking an action a∈A. The state transfer function reflects the possibility of state change after the encryption mechanism is adjusted.

[0074] Reward function r: The reward function r(s, a) is used to evaluate the impact of each encryption mechanism adjustment (transfer from state s to state s′ through action a) on security and performance. Positive rewards indicate that the adjustment of the encryption mechanism has effectively improved security or efficiency, such as successfully resisting attacks or reducing computational overhead. Negative penalties reflect the occurrence of security incidents or the degradation of the performance of the adjusted encryption mechanism. For example, the adjusted encryption strategy is more vulnerable to attacks or causes performance bottlenecks. In addition, the reward function also takes into account balance. For example, an overly strict encryption strategy may reduce the user experience, thereby introducing negative rewards.

[0075] Risk cost function h: The risk cost function h(s) is used to map a state s∈S to a non-negative real value, which is used to evaluate the potential risk cost in that state. For example, some states may correspond to encryption mechanism configurations that are easily broken. This function helps quantify the risk level of different states and serves as a constraint to guide the optimization process.

[0076] Discount factor γ: A discount factor γ close to 1 pays more attention to long-term cumulative rewards, while a γ close to 0 takes short-term rewards into consideration. In the present invention, the discount factor is preferably set to 0.95 or 0.995.

[0077] The goal of the present invention is to maximize the expected sum of discounted rewards while ensuring the overall security of the encryption mechanism optimization process through strategy optimization of a restricted Markov decision process, while ensuring that the total cost (risk) does not exceed a preset threshold.

[0078] Through such an optimization process, the encryption strategy can be continuously adjusted and optimized during the exploration and learning process, so that it gradually moves towards the optimal solution for security and performance.

[0079] Under the premise of ensuring overall safety, the present invention maximizes the expected sum of discounted rewards through restricted Markov decision process strategy optimization, while ensuring that the total cost (risk) does not exceed a preset threshold.

[0080] Step 2: Construct an adversarial autoencoder to evaluate the security of the encryption mechanism, and divide the encryption mechanism into a safe set and a violation set in combination with the risk cost function;

[0081] This paper proposes an evaluation method based on adversarial autoencoders to concretize the risk cost function h(s), which measures the security of the encryption mechanism in the face of tampering. The adversarial autoencoder quantifies the vulnerability of the encryption mechanism by reconstructing the encryption mechanism and detecting the degree of confusion between the reconstruction result and the original data. Its core lies in learning the characteristic distribution of the encryption mechanism and accurately evaluating its security by generating adversarial samples.

[0082] Step 2 specifically includes:

[0083] Step 2-1: Build an adversarial autoencoder and obtain the score of the encryption mechanism through the adversarial autoencoder

[0084] The adversarial autoencoder consists of an autoencoder and an adversarial network, where the autoencoder consists of an encoder and a decoder.

[0085] The encoder maps the input encryption mechanism state into a low-dimensional latent space, while the decoder reconstructs the original encryption mechanism state from these latent spaces. The goal of the reconstruction process is to minimize the mean square error (MSE) between the input state and the reconstructed state, which is used to measure the preservation of the encryption mechanism characteristics during the dimensionality reduction and reconstruction process:

[0086]

[0087] Among them, s i is the original encryption mechanism state, is the reconstructed state output by the decoder, and N is the number of states. The mean square error represents the difference between the reconstructed state and the original state. recon Represents the error loss during the decoder reconstruction process.

[0088] The adversarial network consists of a discriminator that distinguishes the original encryption mechanism state from the encryption mechanism state reconstructed by the autoencoder.

[0089] Adversarial networks achieve this task through two loss functions: true distribution loss and fake distribution loss. The true distribution loss measures the ability of the discriminator to classify real samples as real, while the fake distribution loss measures the ability of the discriminator to classify samples generated by the encoder as fake. These loss functions are calculated using binary cross entropy:

[0090]

[0091] L dis =L dreal +L fake

[0092] Among them, D(s i ) is the discriminator for the real sample s i The output, is the discriminator's output sample to the encoder The score of L, σ is the sigmoid function, log is the natural logarithm, which is used to convert the probability to the logarithmic scale, and N represents the number of states. real represents the true distribution loss, L fake represents the false distribution loss, L dis Represents the error loss in the discriminator's discrimination process.

[0093] In addition, the encoder is used as a generator, and the goal of the generator is to generate encryption mechanism states that the discriminator cannot distinguish, so as to evaluate the tampering resistance and security of the current encryption mechanism.

[0094] The loss function of the generator is defined as:

[0095]

[0096] in, is the discriminator's output sample to the encoder The score of L, σ is the sigmoid function, log is the natural logarithm, which is used to convert the probability to the logarithmic scale, and N represents the number of states. gen Represents the error loss of the generator to generate the encryption mechanism

[0097] Optimizing this generator loss can improve the adversarial resistance of the encoder output, making it more difficult for the discriminator to distinguish as a reconstructed sample, thereby better simulating potential attacks or tampering scenarios.

[0098] Finally, the overall loss function L of the adversarial autoencoder is total The expression is as follows:

[0099] L total =αL recon +βL dis +γL gen

[0100] Among them, α, β and γ are the weight coefficients corresponding to adjusting the autoencoder loss, discriminator loss and generator loss, respectively, which are used to adjust the relative importance of the autoencoder loss, discriminator loss and generator loss in the overall training process.

[0101] In the present invention, the adversarial autoencoder is used as a security assessment tool in the restricted Markov decision process to detect the potential risk cost h of the encryption mechanism during the adjustment process.

[0102] Step 2-2, construct the importance weight formula of the state, which is used to calculate the importance weight of each encryption mechanism;

[0103] However, unlike the previous work on learning to reconstruct the recovery effect through adversarial autoencoders, we note that different encryption mechanism states should be treated differently. Some encryption mechanism states are "critical" and choosing bad adjustment actions will lead to catastrophic consequences. For example, when a certain encryption mechanism state approaches a high-risk area, we should make the network more able to minimize the violation cost.

[0104] The present invention measures the impact of the encryption mechanism state on the reward and uses the following state importance weight formula w(s) to distinguish different encryption mechanism states:

[0105]

[0106] Among them, a1 and a2 represent two actions in the action space A of the restricted Markov decision process, and r(s, a1) and r(s, a2) represent the reward functions of the encryption mechanism state s under actions a1 and a2 respectively.

[0107] By calculating the importance weight w(s), the subsequent policy network can pay more attention to the more critical encryption mechanism status when optimizing.

[0108] Step 2-3, construct a risk cost function, and calculate the risk cost of the encryption mechanism by combining the encryption mechanism's score and importance weight;

[0109] The risk of the encryption mechanism state is quantified by the risk cost function:

[0110] The current original encryption mechanism state is input into the encoder to generate a reconstructed encryption mechanism, and then the discriminator is used to determine whether the reconstructed encryption mechanism can be confused with the original encryption mechanism: if the discriminator believes that the reconstructed encryption mechanism is similar to the original mechanism, it means that the current encryption mechanism has a high risk of being tampered with, which is a high-risk state; if the discriminator believes that the reconstructed encryption mechanism is not similar to the original mechanism, it means that the current encryption mechanism has a low risk of being tampered with, which is a low-risk state. The risk of the encryption mechanism state is quantified by the risk cost function h(s).

[0111] The constructed risk cost function h(s) is as follows:

[0112]

[0113] Among them, D(s) is the score output of the discriminator for the original mechanism s, is the discriminator to reconstruct the encryption mechanism , ε is the risk tolerance factor, max{} is the maximum value operation, and w(s) represents the importance of state s.

[0114] Through this method, the present invention can conduct a comprehensive quantitative evaluation of the security of the encryption mechanism during the optimization process, help optimize decision-making and strategy adjustment, and ensure that the encryption mechanism finally selected has high security and tamper resistance.

[0115] By specifying the risk cost function h(s), the instantaneous security of the encryption mechanism has been clearly focused on. On this basis, the present invention further uses the formalized risk cost function to evaluate the overall feasibility of the encryption mechanism during dynamic adjustment and optimization.

[0116] By introducing the optimal feasibility function, the security of the encryption mechanism under various states and strategy selections can be comprehensively considered to ensure that it always remains secure and effective during the optimization process.

[0117] Step 2-4, combined with the calculation results of the risk cost function, divide the encryption mechanism in the state space S into a security set S s and violation set S v ;

[0118] Define the safety set S s represents all the cryptographic mechanism states that are considered safe. Each state in the safe set represents a well-configured cryptographic mechanism that is not easily attacked or tampered with under the current security situation. Accordingly, all cryptographic mechanism states that do not meet these conditions constitute the violation set S v , violation set S v is the complement of the safe set:

[0119] Ss : ={s∈S:h(s)=0}

[0120] S v : ={s∈S:h(s)>0}

[0121] Among them, h(s) is the risk cost function.

[0122] Step 3, construct the optimal feasibility function to divide the safety set and violation set into feasible set and infeasible set;

[0123] It is not enough to only evaluate the instantaneous security of a single state, especially in the face of ever-changing security threats and policy adjustments. The process of optimizing the encryption mechanism may involve transitions between multiple states, and the action chosen at each step may affect the overall security.

[0124] Therefore, in order to evaluate whether the encryption mechanism is safe during the entire adjustment and optimization process, the present invention defines the optimal feasibility function H*(s) to find the strategy that can maximize security during the optimization process. The optimal feasibility estimation function H* is used to calculate the probability of violation when starting from an initial state, through a series of strategies and action selections, along the path determined by the state transition model, and finally reaching the terminal state.

[0125]

[0126] in, is the instantaneous violation indication function. When the state s t When it is not in the security set, its value is 1; if it is in the security set, its value is 0. τ~π*, P(s) means starting from state s, using the optimal security strategy π* to sample trajectory τ under the state transition function P, For violation set, represents the expected discounted cost.

[0127] By taking the expectation of all possible trajectories, we can calculate the probability of violation when the encryption mechanism starts from the initial state s0 and goes through the adjustment process under the optimal security policy π*.

[0128] In the present invention, the feasible set of the optimal security strategy π* is It can be described as a set of states that ensure that there is no entry into the unsafe set (violation set) on the entire path during all possible state transitions of the encryption mechanism, and the infeasible set is a feasible set The complement of :

[0129]

[0130] Among them, H*(s) represents the optimal feasibility function.

[0131] By evaluating the feasibility of the encryption mechanism, the security set and the violation set are further divided into a feasible set and an infeasible set. On this basis, the present invention proposes a security policy optimization method for feasibility evaluation of encryption mechanisms, aiming to achieve efficient tuning and long-term security assurance of encryption mechanisms in dynamic environments. The method designs different optimization strategies for encryption mechanisms in feasible sets and infeasible sets, respectively, so as to ensure the security and effectiveness of the entire optimization process. The method of the present invention first divides the encryption mechanism state set into feasible sets by using the feasibility estimation function H*(s). and infeasible set

[0132] Step 4: Based on the feasibility of the encryption mechanism, the security policy is used to optimize the encryption mechanism of the infeasible centralization and dynamically adjust the security policy.

[0133] In the present invention, by using the optimal feasibility function to evaluate and guide the optimization process of the encryption mechanism, it can be ensured that the policy adjustment from the initial state to the terminal state is always within the safe and feasible range. Through this method, the present invention can optimize the encryption mechanism in the feasible set more accurately, and for the encryption mechanism not in the feasible set, the policy is adjusted to return to the safe range, thereby avoiding potential security risks in the encryption mechanism optimization process. Such optimization not only improves the security of the encryption mechanism, but also enhances its overall efficiency. The specific implementation steps are as follows:

[0134] For the encryption mechanism state in the infeasible set, the focus of optimization is to quickly guide it back to the state of the feasible set, thereby reducing security risks. This is achieved by minimizing the maximum possible violation cost in the future. The process involves predicting future states and evaluating the violation risks that may be caused by different action choices. The optimization objective function is as follows:

[0135]

[0136] Among them, h(s) is the risk cost function, s t is the tth state.

[0137] In this way, policy adjustments can be made quickly in high-risk situations to avoid unacceptable security incidents.

[0138] On the other hand, for the encryption mechanism state within the feasible set, the optimization goal is to further optimize its performance and effectiveness under the premise of acceptable security. This is usually achieved by maximizing the expected reward, that is, improving the efficiency of the encryption mechanism as much as possible while ensuring security. Specifically, the present invention uses a constrained optimization problem solving strategy based on the Lagrangian method to achieve optimal adjustment. The core of this process is to find a balance between security (constraints) and rewards (effectiveness). Specifically, through the Lagrangian multiplier method, an optimization model is constructed to convert the security constraints into penalty terms in the optimization problem, and maximize the rewards while ensuring basic security. The objective function of the optimization process is:

[0139]

[0140] Where λ is the Lagrange multiplier, c is the threshold of the safety constraint, γ is the discount factor, and r(s t , a t ) is the reward function, h(s t ) is the risk cost function.

[0141] In the present invention, the effective use of the feasibility estimation function H*(s) to determine whether a state is in a feasible set depends on continuous interaction with the environment. In the absence of reliable prior knowledge, the feasibility estimation function defined recursively can gradually learn and predict the safety of the state, thereby improving the accuracy and efficiency of feasibility judgment. The recursive definition is as follows:

[0142]

[0143] Where s'~π, P(s) is a sample obtained by sampling the successor state (i.e., s'~P(·|s, a~π(·|s))), and it is expected to be taken from all possible successor states, This recursive definition allows the feasibility of gradually learning and accurately predicting states through interaction with the environment without complete prior knowledge.

[0144] Therefore, the strategy optimization process of the entire encryption mechanism is as follows:

[0145]

[0146] Where d0 is the initial state distribution, H*(s) represents the optimal feasibility function, λ is the Lagrange multiplier, c is the threshold of the safety constraint, and γ is the discount factor.

[0147] Further preferably, the policy optimization process of the present invention is combined with a modern deep reinforcement learning framework, in particular, with an actor-critic method, such as Figure 3As shown, the strategy optimization process is as follows:

[0148]

[0149] Where π represents the strategy, a~π(·|s) represents the sampling action a according to the strategy, Q(s, a) is the state-action value function, and Q c (s, a) is the state-action cost value function.

[0150] Further preferably, through iterative calculation and updating of policy network parameters, the optimal policy can be gradually approached in a complex multi-state, multi-action space, thereby realizing the update of the policy optimization process:

[0151] State-action value Q network parameter ω, state-action cost value Q c The network parameters p, policy network parameters θ and Lagrange multipliers λ are updated as follows:

[0152]

[0153] in, represents the gradient, ξ1, ξ2, ξ3 and ξ4 are the learning rates of the network, which can be dynamically adjusted according to the effect of the current policy execution and the expected security requirements to ensure the adaptability and flexibility of the policy. Through this optimization method, the present invention can dynamically adjust and optimize the encryption mechanism strategy, thereby achieving effective improvement of the initial encryption mechanism. Ultimately, the generated optimization strategy can be used to adjust the initial encryption mechanism, thereby generating a new, more secure and efficient encryption mechanism to adapt to the ever-changing security environment.

[0154] like Figure 4 As shown, the present invention also proposes a dynamic optimization system for encryption mechanism based on restricted Markov decision process, which is used to implement the above-mentioned dynamic optimization method for encryption mechanism based on restricted Markov decision process, and the system includes: a restricted Markov decision process building module, a primary partitioning module, a secondary partitioning module and an optimization adjustment module;

[0155] The restricted Markov decision process building module is used to build a restricted Markov decision process based on the encryption mechanism to be optimized, and the restricted Markov decision process includes a risk cost function;

[0156] The initial partitioning module evaluates the security of the encryption mechanism by constructing an adversarial autoencoder and divides the encryption mechanism into a secure set and a violation set in combination with a risk cost function;

[0157] The secondary partitioning module divides the safety set and violation set into feasible set and infeasible set by constructing the optimal feasibility function;

[0158] The optimization and adjustment module optimizes the encryption mechanism of the infeasible concentration based on the security policy of the feasibility of the encryption mechanism and dynamically adjusts the security policy.

[0159] The beneficial effect of the present invention is that, compared with the prior art, the present invention constructs an optimal feasibility function to evaluate the feasibility of the encryption mechanism adjustment sequence from the initial state to the terminal state, ensuring that the overall security and stability are always maintained during the optimization process of the encryption mechanism; at the same time, an optimization strategy adjustment method is designed to enable the encryption mechanism to gradually improve its security and performance in the process of continuous strategy optimization, and ultimately achieve self-learning and dynamic adjustment of the encryption mechanism optimization strategy.

[0160] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0161] Computer readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. Computer readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive lists) of computer readable storage medium include: portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), static random access memories (SRAM), portable compact disk read-only memories (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or convex structures in grooves on which instructions are stored, and any suitable combination thereof. Computer readable storage medium used here is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (e.g., light pulses by optical fiber cables), or electrical signals transmitted by wires.

[0162] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0163] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages-such as Smalltalk, C++, etc., and conventional procedural programming languages-such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network-including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may be personalized by utilizing the state information of a computer-readable program instruction, and the electronic circuit may execute a computer-readable program instruction, thereby realizing various aspects of the present disclosure.

[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for dynamic optimization of encryption mechanism based on restricted Markov decision process, characterized in that: The steps include: Based on the encryption mechanism to be optimized, a restricted Markov decision process is constructed, and the restricted Markov decision process includes a risk cost function; Construct an adversarial autoencoder to evaluate the security of the encryption mechanism, and divide the encryption mechanism into a safe set and a violation set in combination with the risk cost function; Construct the optimal feasibility function to divide the safety set and violation set into feasible set and infeasible set; The security strategy based on the feasibility of the encryption mechanism optimizes the encryption mechanism of the infeasible concentration and dynamically adjusts the security strategy.

2. The method for dynamic optimization of encryption mechanism based on restricted Markov decision process according to claim 1, characterized in that: The restricted Markov decision process includes state space, action space, state transition function, reward function, risk cost function and discount factor; Among them, the state space S represents the current encryption mechanism state, covering various parameters of the encryption mechanism state and security situation indicators; The action space A represents the possible policy adjustment operations that the encryption mechanism may take in each state; The state transition function P(s′|s, a) describes the probability of transitioning from the current encryption mechanism state s∈S to the next encryption mechanism state s′∈S after taking action a∈A; The reward function r(s, a) is used to evaluate the impact of each encryption mechanism state transition on security and performance; The security loss function h is used to map the encryption mechanism state s∈S to a non-negative real value, which is used to evaluate the potential risk cost in the encryption mechanism state.

3. The method for dynamic optimization of encryption mechanism based on restricted Markov decision process according to claim 1, characterized in that: The construction of the adversarial automatic encoder specifically includes: The adversarial autoencoder includes an autoencoder and an adversarial network, wherein the autoencoder includes an encoder and a decoder, the encoder is used to map the input encryption mechanism state into a low-dimensional latent space, and the decoder is used to obtain the reconstructed encryption mechanism state from the low-dimensional latent space; the adversarial network includes a discriminator, which is used to distinguish the original encryption mechanism state from the encryption mechanism state reconstructed by the autoencoder.

4. The method for dynamic optimization of encryption mechanism based on restricted Markov decision process according to claim 3 is characterized in that: The risk cost function specifically includes: The risk cost function is used to calculate the risk cost h(s) of the encryption mechanism s: Where D(s) is the score output of the discriminator for the encryption mechanism state s, is the discriminator to reconstruct the encryption mechanism , ε is the risk tolerance factor, max{} is the maximum value operation, and w(s) represents the importance of the encryption mechanism state s.

5. The method for dynamic optimization of encryption mechanism based on restricted Markov decision process according to claim 4, characterized in that: The combined risk cost function divides the encryption mechanism into a safe set and a violation set, specifically including: According to the risk cost function h(s), the encryption mechanism s is judged to belong to the security set or the violation set: S s :={s∈S:h(s)=0} S v :={s∈S:h(s)>0} Where S represents the state space of the restricted Markov decision process, S s represents the safe set, S v Represents the offending set.

6. The method for dynamic optimization of encryption mechanism based on restricted Markov decision process according to claim 5, characterized in that: By measuring the impact of the encryption mechanism state on the reward, the importance weight w(s) of the encryption mechanism state s is calculated as follows: Among them, a1 and a2 represent two actions in the action space A of the restricted Markov decision process, and r(s, a1) and r(s, a2) represent the reward functions of the encryption mechanism state s under actions a1 and a2 respectively.

7. The method for dynamic optimization of encryption mechanism based on restricted Markov decision process according to claim 1, characterized in that: The optimal feasibility function H is constructed * (s), including: in, is the instantaneous violation indicator function, when the encryption mechanism state s at time t t When it is not in the security set, its value is 1; when the encryption mechanism state s at time t t In the safe concentration, its value is 0; T~π * , P(s) means starting from the encryption mechanism state s, using the optimal security strategy π * Sampling trajectory τ under the state transfer function P, For violation set.

8. The method for dynamic optimization of encryption mechanism based on restricted Markov decision process according to claim 7, characterized in that: The division of the safety set and the violation set into a feasible set and an infeasible set specifically includes: According to the calculated optimal feasibility function H * The value of (s) divides the safe set and the violation set into feasible set and infeasible set: in, represents the feasible set, Represents an infeasible set.

9. The method for dynamic optimization of encryption mechanism based on restricted Markov decision process according to claim 1, characterized in that: The security strategy based on the feasibility of the encryption mechanism optimizes the encryption mechanism of the infeasible centralization, specifically including: Combined with the actor-critic algorithm, the encryption mechanism in the infeasible set is guided to return to the feasible set by minimizing the maximum possible violation cost in the future, and the encryption mechanism in the feasible set is optimized by maximizing the expected reward. The resulting security strategy process is as follows: Where π represents the strategy, a~π(·|s) represents the sampling action a according to the strategy, Q(s, a) is the state-action value function, and Q c (s, a) is the state-action cost value function, H * (s) is the optimal feasibility function, λ is the Lagrange multiplier, d0 is the initial encryption mechanism state distribution, s~d0 represents the encryption mechanism state sampled from the initial encryption mechanism state distribution, and c is the threshold constant.

10. The method for dynamic optimization of encryption mechanism based on restricted Markov decision process according to claim 1, characterized in that: Dynamically adjust security policies, including: The network parameters ω of the state-action value function Q(s, a), the state-action cost value function Q c The network parameter p of (s, a), the parameter θ of the policy network π, and the Lagrange multiplier λ are dynamically updated as follows: Among them, s t 、s t+1 Respectively represent the encryption mechanism status at time t and time t+1, a t 、a t+1 Respectively represent the actions on the encryption mechanism state at time t and time t+1, represents the gradient, They represent the gradient of parameters ω, p, and λ respectively, (1, ζ2, ζ3, and ζ4 are the learning rates of the network, π θ represents a policy with parameter θ.

11. A dynamic optimization system for encryption mechanism based on restricted Markov decision process, used to implement the dynamic optimization method for encryption mechanism based on restricted Markov decision process as claimed in any one of claims 1 to 10, characterized in that: include: Restricted Markov decision process building module, primary partitioning module, secondary partitioning module and optimization adjustment module; The restricted Markov decision process building module is used to build a restricted Markov decision process based on the encryption mechanism to be optimized, and the restricted Markov decision process includes a risk cost function; The initial partitioning module evaluates the security of the encryption mechanism by constructing an adversarial autoencoder and divides the encryption mechanism into a secure set and a violation set in combination with a risk cost function; The secondary partitioning module divides the safety set and violation set into feasible set and infeasible set by constructing the optimal feasibility function; The optimization and adjustment module optimizes the encryption mechanism of the infeasible concentration based on the security policy of the feasibility of the encryption mechanism and dynamically adjusts the security policy.

12. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.