Strategy optimization method of high-security wireless communication system based on hierarchical game
By optimizing the communication delay allocation ratio and security strategy of the wireless communication system based on a hierarchical game method, the problem that the wireless communication system cannot balance the communication delay and security protection strength when defending against security threats is solved, and a dynamic balance of high security and high efficiency is achieved in the wireless communication system.
Patent Information
- Application Number
- CN202510827051.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-26
AI Technical Summary
Existing wireless communication systems are unable to effectively strike a dynamic balance between communication latency and security protection strength when defending against security threats, resulting in wasted communication resources or rigid security policies, making it difficult to adapt to the ever-changing network environment and rapidly evolving attack situations.
A hierarchical game-based approach is adopted, with the access network and core network as independent game participants. By constructing a total variable game model, with the goal of optimizing the overall communication security benefits, the communication delay allocation ratio and security strategy are iteratively optimized, and a continuously iterative hierarchical game global optimization interaction mechanism is established to achieve the optimal dynamic allocation of delay resources.
It realizes the dynamic game balance between the access network and the core network, achieves the overall coordinated optimization of the communication efficiency and security assurance capabilities of the wireless communication system, and ensures the intelligent and stable operation of the highly secure wireless communication system in multiple scenarios.
Smart Images

Figure CN120711399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication security technology, and in particular to a strategy optimization method for a high-security wireless communication system based on hierarchical game. Background Art
[0002] With the rapid development of 5G, satellite communications, and future mobile communication networks, wireless communication systems in key areas such as military, finance, healthcare, and government confidential communications are facing increasingly severe and complex security threats, such as eavesdropping, data tampering, and denial of service (DoS) attacks.
[0003] Currently, commonly used methods to defend against security threats include encryption enhancement, channel optimization, network protection and other security strategies.
[0004] However, these methods fail to effectively balance the dynamic balance between communication latency, bit error rate, and security protection strength when defending against security threats, resulting in wasted communication resources or rigid security policies, making it difficult to adapt to the ever-changing network environment and rapidly evolving attack situations. Summary of the Invention
[0005] The present invention provides a policy optimization method for a high-security wireless communication system based on hierarchical game, which is used to solve the defect in the existing technology that the wireless communication system cannot simultaneously take into account the dynamic balance between communication delay and security protection strength when defending against security threats, and realizes a communication security policy optimization method based on hierarchical game.
[0006] The present invention provides a strategy optimization method for a high-security wireless communication system based on hierarchical game, comprising: S10. Determine a current communication delay division ratio of the wireless communication system, a first current communication security policy of the access network, and a second current communication security policy of the core network; S20. With the goal of optimizing the overall communication security benefits of the access network and the core network, iteratively optimize the current communication delay division ratio to obtain an optimized communication delay division ratio; S30. Optimize the first current communication security policy and the second current communication security policy under the optimized communication delay division ratio to obtain a first optimized communication security policy for the access network and a second optimized communication security policy for the core network; S40, determining a new current communication delay division ratio, a new first current communication security policy, and a new second current communication security policy based on the optimized communication delay division ratio, the first optimized communication security policy, and the second optimized communication security policy; S50. Iteratively execute steps S20 to S40 to obtain a global optimal communication security policy for the wireless communication system; the global optimal communication security policy includes a first optimal communication security policy for the access network and a second optimal communication security policy for the core network.
[0007] According to a policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention, step S20 includes: Under the first KKT condition, with the goal of optimizing the overall communication security benefit of the access network and the core network, iteratively optimize the current communication delay partition ratio based on a gradient descent method and a first Lagrangian objective function to obtain the optimized communication delay partition ratio; The first Lagrangian objective function includes an external joint utility function and an upper limit constraint on communication delay allocation; The external joint utility function includes an access network utility function and a core network utility function; the access network utility function is determined based on a communication policy benefit and a communication policy overhead; the core network utility function is determined based on a security policy benefit and a security policy overhead; the communication policy benefit and the communication policy overhead are determined based on a first current communication security policy and a current communication delay division ratio; the security policy benefit and the security policy overhead are determined based on a second current communication security policy and a current communication delay division ratio; The communication delay allocation upper limit constraint is determined based on the current total budget of end-to-end delay, the adjustable upper bound of the total budget of end-to-end communication delay, and the resource sensitivity of the adjustable upper bound of the total budget of end-to-end communication delay; the Lagrangian multiplier of the first Lagrangian objective function is the resource sensitivity.
[0008] According to a policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention, step S30 includes: S31, with the goal of optimizing the cooperative joint utility between the first communication decision-maker and the first security decision-maker of the access network under the optimized communication delay division ratio, iteratively optimize the first current communication security policy to obtain a first optimized communication security policy; S32. With the goal of optimizing the competitive joint utility of the second communication decision-maker and the second security decision-maker of the core network under the optimized communication delay division ratio, the second current communication security policy is iteratively solved to obtain a second optimized communication security policy.
[0009] According to a policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention, the first current communication security policy includes the communication policies of the first communication decision-maker and the security policies of the first security decision-maker; the cooperative joint utility includes the utility of the communication decision-maker and the utility of the security decision-maker; step S31 includes: Under a second KKT condition, with the goal of optimizing the cooperative joint utility of the access network, the first current communication security policy is iteratively optimized based on a policy update rule and a second Lagrangian objective function to obtain the first optimized communication security policy; the policy update rule includes sub-update rules for each communication policy and each security policy in the first current communication security policy; The second Lagrangian objective function includes a cooperative joint utility function, an access network communication delay budget constraint, a bandwidth resource constraint, and a security constraint; The cooperative joint utility function is determined based on the communication decision-maker utility function and the security decision-maker utility function; the communication decision-maker utility function is determined based on the policy benefits of each communication policy of the first communication decision-maker and the delay overhead of each security policy of the first security decision-maker; the security decision-maker utility function is determined based on the anti-attack capability of each security policy of the first security decision-maker and the security synergy benefit of each communication policy of the first communication decision-maker; The access network communication delay budget constraint is determined based on the optimized communication delay division ratio and the access network delay budget sensitivity; the bandwidth resource constraint is determined based on the access network allocated bandwidth, the total system available bandwidth and the bandwidth resource sensitivity in the first current communication strategy; the security constraint is determined based on the minimum access network security utility, the security decision-maker utility and the security sensitivity; the Lagrangian multiplier of the second Lagrangian objective function includes the access network delay budget sensitivity, the bandwidth resource sensitivity and the security sensitivity.
[0010] According to a policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention, the method takes the optimization of the cooperative joint utility of the access network as the goal, iteratively optimizes the first current communication security policy based on the policy update rule and the second Lagrangian objective function to obtain the first optimized communication security policy, including: constructing a first input vector based on the first current communication security policy; Inputting the first input vector into a first policy update model to obtain a first optimized communication security policy output by the first policy update model; the first policy update model is implemented based on a deep Q network; The state space of the first policy update model is determined based on the communication strategies and security strategies in the first current communication security strategy; the action space of the first policy update model is determined based on the policy update rule; and the reward function of the first policy update model is determined based on the cooperative joint utility function in the second Lagrangian objective function.
[0011] According to a policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention, when the second current communication security policy belongs to a discrete policy, step S32 includes: Based on the second current communication security policy, defining a communication hybrid policy set of a second communication decision-maker of the core network and a security hybrid policy set of a second security decision-maker of the core network; defining a first hybrid strategy probability distribution variable corresponding to the communication hybrid strategy set and a second hybrid strategy probability distribution variable corresponding to the security hybrid strategy set; Constructing a game matrix; each element in the game matrix is a loss function value obtained by substituting a communication hybrid strategy from the communication hybrid strategy set and a security hybrid strategy from the security hybrid strategy set into a pre-constructed communication party loss function; the communication party loss function is determined based on the core network end-to-end delay and the security policy communication overhead; Based on the optimized communication delay partition ratio, the game matrix, and the first mixed strategy probability distribution variable, with the goal of minimizing the upper limit of communication expected loss, constructing the original linear programming model of the second communication decision-maker; Based on the optimized communication delay partition ratio, the game matrix, and the second mixed strategy probability distribution variable, a dual linear programming model of the second security decision-maker is constructed with the goal of maximizing the lower bound of the expected security benefit; Calling a standard linear programming solver to jointly solve the original linear programming model and the dual linear programming model to obtain an optimal solution pair of the hybrid strategy probability distribution with the optimal competitive joint utility; The second optimized communication security strategy is determined based on the optimal solution pair of the hybrid strategy probability distribution.
[0012] According to a policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention, when the second current communication security policy is a continuous policy, step S32 includes: constructing a second input vector based on the communication policy in the second current communication security policy; Inputting the second input vector into a second policy update model to obtain an optimized communication policy output by the second policy update model; the second policy update model is implemented based on a deep Q network; constructing a third input vector based on the security policy in the second current communication security policy; Inputting the third input vector into a third policy update model to obtain an optimized security policy output by the third policy update model; the third policy update model is implemented based on a deep Q network; Based on the optimized communication strategy and the optimized security strategy, obtaining the second optimized communication security strategy; The state spaces of the second policy update model and the third policy update model are both determined based on each communication policy and each security policy in the second current communication security policy; The action space of the second policy update model is determined based on the communication policy in the second current communication security policy; the reward function of the second policy update model is determined based on the core network end-to-end delay and the security policy communication overhead; The action space of the third policy update model is determined based on the security policy in the second current communication security policy; the reward function of the third policy update model is determined by the security policy benefit and the security policy encryption processing delay.
[0013] According to a policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention, the iterative execution of steps S20 to S40 to obtain a global optimal communication security policy for the wireless communication system includes: Iteratively executing steps S20 to S40 until the optimized communication delay division ratio converges, or until the first optimized communication security strategy and the second optimized communication security strategy are stable, thereby obtaining the first optimal communication security strategy and the second optimal communication security strategy; The global optimal communication security policy is determined based on the first optimal communication security policy and the second optimal communication security policy.
[0014] According to a policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention, the security strength of the global optimal communication security policy is greater than the minimum security strength of the mobile communication service; the minimum security strength is determined based on the following method; Determine the current wireless channel environment threat index based on the abnormal behavior index of the current access requesting terminal and the frequency of unsuccessful matching between the access point and the terminal security mechanism per unit time; determining a risk index for the mobile communication service based on the sensitivity of the mobile communication service to security risks and the threat index of the current wireless channel environment; The minimum security strength is determined based on the risk index of the mobile communication service, the service fixed security strength level, and the maximum risk level that can be coped with by the service fixed security strength level.
[0015] According to a policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention, S60, based on the global optimal communication security policy, determines the target security function module to be called; calls the target security function module from the scalable security function resource pool to realize secure communication of the wireless communication system.
[0016] The present invention provides a strategy optimization method for a high-security wireless communication system based on hierarchical game. By constructing a total variable game model, the access network and the core network of the wireless communication system are regarded as independent game participants. Based on the Nash equilibrium principle, the communication delay division ratio is first iteratively optimized with the goal of optimizing the overall communication security benefit. Then, the communication security strategy of the access network and the communication security strategy of the core network are optimized under the optimized communication delay division ratio. A global optimization interaction mechanism of continuously iterative hierarchical game is established to solve the optimal dynamic allocation strategy of delay resources. The optimal global communication security strategy of the wireless communication system is solved, so that a dynamic game equilibrium is achieved between the access network and the core network, thereby achieving the overall synergy optimization between the communication efficiency and security assurance capability of the wireless communication system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 It is a flow chart of a strategy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention.
[0019] Figure 2 It is a structural diagram of a strategy optimization device for a high-security wireless communication system based on hierarchical game provided by the present invention.
[0020] Figure 3 It is a schematic diagram of the hierarchical game mechanism architecture provided by the present invention. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0022] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or apparatus. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0023] The terms "first," "second," and the like in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that embodiments of the present invention can be implemented in orders other than those illustrated or described herein. Furthermore, the objects distinguished by "first," "second," and the like generally refer to a class of objects, and do not limit the number of objects. For example, the first object may be one or more.
[0024] The following combination Figure 1-Figure 3 The present invention describes a strategy optimization method for a high-security wireless communication system based on hierarchical game.
[0025] Figure 1 Schematic diagram of the process of the strategy optimization method of the high-security wireless communication system based on hierarchical game provided by the present invention. Figure 1 As shown, the strategy optimization method for a high-security wireless communication system based on hierarchical game includes but is not limited to steps S10 to S50.
[0026] It should be noted that the executor of the policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention is the corresponding policy optimization device for a high-security wireless communication system based on hierarchical game, which can specifically be a server or computer equipment, such as a mobile phone, tablet computer, laptop computer, PDA, vehicle-mounted electronic equipment, wearable device, ultra-mobile personal computer (UMPC), netbook or personal digital assistant (PDA), etc.
[0027] Step S10: Determine the current communication delay division ratio of the wireless communication system, the first current communication security policy of the access network, and the second current communication security policy of the core network.
[0028] Wireless communication systems include access networks and core network In a highly secure wireless communication system, the access network and the core network compete for the allocation of end-to-end tolerable delay. The access network seeks a larger delay budget to optimize communication decisions, while the core network seeks a more ample delay budget to increase the delay resources required for security policy implementation and strengthen data security protection. Therefore, the access network and the core network are participants in the external competitive game, and the access network and the core network must conduct dynamic game optimization under the total delay constraint. In other words, the access network and the core network conduct an external competitive game based on the total variable game framework.
[0029] The current total budget of end-to-end delay determined in the wireless communication system Next, the communication delay division ratio is introduced , the optimization goal of the access network is to maximize the end-to-end tolerable delay that can be allocated to it The core network's optimization goal is to maximize its own allocable end-to-end tolerable delay. .
[0030] Based on the external competition game between the access network and the core network, the communication delay allocation ratio is iteratively optimized, which enables the access network to maximize the communication performance within its own allocable end-to-end tolerable delay, and the core network to maximize the security performance within its own allocable end-to-end tolerable delay.
[0031] in, Communication delay division ratio As the core decision variable in the external game between the access network and the core network in the wireless communication system, it can control the ratio of the delay resources that can be allocated between the two in the communication system.
[0032] In addition, the external game between the access network and the core network needs to meet the adjustable upper bound of the total budget of end-to-end communication delay The constraint is the expression that satisfies the following upper limit constraint of communication delay allocation: (1) Both the access network and the core network include communication decision-makers and security decision-makers.
[0033] The communication decision-maker included in the access network is the first communication decision-maker The security decision-maker included in the access network is the first security decision-maker , the first communication decision maker and the first security decision maker are game participants in the cooperation game within the access network.
[0034] The communication decision-maker included in the core network is the second communication decision-maker , core network The included security decision maker is the second security decision maker ,The second communication decision maker and the second security decision maker are the game participants in the ,competitive game within the core network.
[0035] When defending against security threats, the first communication decision-maker adopts several communication strategies, and the first security decision-maker adopts several security strategies. Together, these strategies form the access network's communication security strategy, also known as the first communication security strategy. The second communication decision-maker adopts several communication strategies, and the second security decision-maker adopts several security strategies. Together, these strategies form the core network's communication security strategy, also known as the second communication security strategy. Based on the first and second communication security strategies, the overall communication security benefits of the access and core networks in the wireless communication system can be determined.
[0036] Specifically, the communication delay partition ratio of the wireless communication system, the first communication security policy of the access network, and the second communication security policy of the core network are initialized, thereby obtaining a current communication delay partition ratio, a first current communication security policy, and a second current communication security policy. That is, the current communication delay partition ratio is the initialized communication delay partition ratio, the initialized first communication security policy, and the initialized second communication security policy.
[0037] Alternatively, the communication delay division ratio, the first communication security policy of the access network and the second communication security policy of the core network of any iteration during the iterative optimization process of the communication delay division ratio are determined as the current communication delay division ratio, the first current communication security policy and the second current communication security policy.
[0038] Step S20: With the goal of optimizing the overall communication security benefit of the access network and the core network, the current communication delay division ratio is iteratively optimized to obtain an optimized communication delay division ratio.
[0039] Specifically, in the external competition game between the access network and the core network of the wireless communication system, with the goal of optimizing the overall communication security benefits of the access network and the core network, the current communication delay division ratio is iteratively optimized multiple times until the iterative termination conditions of the external competition game are met, and the optimized communication delay division ratio is obtained.
[0040] Step S30: Optimize the first current communication security policy and the second current communication security policy under the optimized communication delay division ratio to obtain a first optimized communication security policy of the access network and a second optimized communication security policy of the core network.
[0041] Specifically, after iteratively optimizing the current communication delay division ratio through the external competitive game between the access network and the core network, the first current communication security strategy and the second current communication security strategy are further optimized based on the internal game between the access network and the core network at a given optimized communication delay division ratio to obtain the first optimized communication security strategy of the access network and the second optimized communication security strategy of the core network.
[0042] For example, under the optimized communication delay division ratio, based on the internal cooperative game between the first communication decision-maker and the first security decision-maker within the access network, the first current communication security policy is iteratively optimized to obtain the first optimized communication security policy of the access network; and, under the optimized communication delay division ratio, based on the internal competitive game between the second communication decision-maker and the second security decision-maker within the core network, the second current communication security policy is iteratively optimized to obtain the second optimized communication security policy of the core network.
[0043] Step S40: Determine a new current communication delay division ratio, a new first current communication security policy, and a new second current communication security policy based on the optimized communication delay division ratio, the first optimized communication security policy, and the second optimized communication security policy.
[0044] Specifically, the optimized communication delay division ratio after the external competition game between the access network and the core network is optimized, and the first optimized communication security strategy and the second optimized communication security strategy after the internal game between the access network and the core network are used as the new current communication delay division ratio, the new first current communication security strategy and the new second current communication security strategy.
[0045] S50 , iteratively executing steps S20 to S40 to obtain a global optimal communication security policy for the wireless communication system.
[0046] The global optimal communication security policy includes a first optimal communication security policy of the access network and a second optimal communication security policy of the core network.
[0047] Specifically, after obtaining a new current communication delay partition ratio, a new first current communication security policy, and a new second current communication security policy, steps S20 and S40 are iteratively executed, i.e., a global game iterative optimization is performed on the communication delay partition ratio, the first communication security policy, and the second communication security policy until an iterative termination condition is satisfied, thereby obtaining a global optimal communication security policy consisting of the first optimal communication security policy of the access network and the second optimal communication security policy of the core network. Furthermore, in each global game optimization, the current communication delay partition ratio is first iteratively optimized based on the external competitive game between the access network and the core network to obtain an optimized communication delay partition ratio. Then, under the optimized communication delay partition ratio, the first current communication security policy and the second current communication security policy are optimized based on the internal game between the access network and the core network to obtain the first optimal communication security policy and the second optimal communication security policy of the core network.
[0048] It should be noted that to improve model accuracy and solution stability, this paper normalizes all utility function terms related to communication benefits and resource consumption to unity. The resulting results are used as dimensionless indicators in the optimization modeling, facilitating efficient convergence of the subsequent solver while maintaining the differentiability and convexity of the objective function.
[0049] The present invention provides a strategy optimization method for a high-security wireless communication system based on hierarchical game. By constructing a total variable game model, the access network and the core network of the wireless communication system are regarded as independent game participants. Based on the Nash equilibrium principle, the communication delay division ratio is first iteratively optimized with the goal of optimizing the overall communication security benefit. Then, the communication security strategy of the access network and the communication security strategy of the core network are optimized under the optimized communication delay division ratio. A global optimization interaction mechanism of continuously iterative hierarchical game is established to solve the optimal dynamic allocation strategy of delay resources. The optimal global communication security strategy of the wireless communication system is solved, so that a dynamic game equilibrium is achieved between the access network and the core network, thereby achieving the overall synergy optimization between the communication efficiency and security assurance capability of the wireless communication system.
[0050] Based on the above embodiment, as an optional embodiment, step S20 includes: Under the first KKT condition, with the goal of optimizing the overall communication security benefit of the access network and the core network, iteratively optimize the current communication delay partition ratio based on a gradient descent method and a first Lagrangian objective function to obtain the optimized communication delay partition ratio; The first Lagrangian objective function includes an external joint utility function and an upper limit constraint on communication delay allocation; The external joint utility function includes an access network utility function and a core network utility function; the access network utility function is determined based on a communication policy benefit and a communication policy overhead; the core network utility function is determined based on a security policy benefit and a security policy overhead; the communication policy benefit and the communication policy overhead are determined based on a first current communication security policy and a current communication delay division ratio; the security policy benefit and the security policy overhead are determined based on a second current communication security policy and a current communication delay division ratio; The communication delay allocation upper limit constraint is determined based on the current total budget of end-to-end delay, the adjustable upper bound of the total budget of end-to-end communication delay, and the resource sensitivity of the adjustable upper bound of the total budget of end-to-end communication delay; the Lagrangian multiplier of the first Lagrangian objective function is the resource sensitivity.
[0051] Specifically, before conducting the external competition game between the access network and the core network, it is necessary to pre-establish the Lagrangian objective function for this external competition game, also known as the first Lagrangian objective function. Furthermore, to solve the communication delay partition ratio under inequality constraints, the KKT optimality condition is introduced. This KKT optimality condition provides the necessary optimal solution criteria for constrained nonlinear optimization problems, ensuring convergence and global equilibrium of the solution. In other words, before conducting the external competition game between the access network and the core network, it is necessary to pre-establish the KKT conditions for the external competition game constraints, also known as the first KKT conditions.
[0052] When constructing the first Lagrangian objective function, in order to quantitatively analyze the resource game relationship between the access network and the core network under the end-to-end delay budget, the access network utility function is first constructed. and core network utility function , and based on this, define the external joint utility function of the wireless communication system as a whole , to measure the overall communication security benefits of the access network and core network under different delay distributions.
[0053] In one embodiment, the expression of the access network utility function is as follows: ;(2) in, For access network utility; Benefits of communication strategy for access network; It is the communication strategy overhead of the access network.
[0054] Optionally, the communication strategy benefit of the access network is that the access network obtains available latency The communication benefits brought about by the increase in data throughput based on the first current communication security policy; the communication policy overhead of the access network is the resource overhead generated by the access network in the process of improving the communication capability based on the first current communication security policy, including but not limited to the additional overhead caused by access point selection optimization, power improvement and larger bandwidth resource allocation.
[0055] The optimization goal of the access network is to maximize the communication performance under the end-to-end tolerable delay that can be allocated by itself, that is, .
[0056] In one embodiment, the expression of the core network utility function is as follows: ;(3) in, For core network utility; Benefits for the core network's security strategy; It is the security policy overhead of the core network.
[0057] Optionally, the core network security policy benefit is that the core network obtains available latency The communication security benefits obtained by the wireless communication system through the second current communication security strategy such as encryption, authentication, and integrity verification; the security policy overhead of the core network is the additional resource overhead such as computing, storage, and keys consumed by the second current communication security strategy of the second security decision-maker.
[0058] The optimization goal of the core network is to maximize security performance under the end-to-end tolerable delay that can be allocated by itself, that is, .
[0059] Furthermore, the expression of the external joint utility function defined according to the access network utility function and the core network utility function is as follows: ;(4) in, is the external joint utility function; is the access network utility function; is the core network utility function.
[0060] The optimization goal of the wireless communication system is to maximize the external joint utility function, that is, .
[0061] When constructing the first Lagrangian objective function, in order to achieve the optimization of the external joint utility function under the constraints, it is also necessary to introduce the upper limit constraint of the communication delay allocation to limit the total communication delay budget of the wireless communication system to not exceed the upper bound.
[0062] Optionally, the expression of the upper limit constraint of the communication delay allocation is as follows: ;(5) in, is the current total budget for end-to-end delay; The adjustable upper bound of the total budget for end-to-end communication delay; Allocate a ratio for communication delay.
[0063] Based on the external joint utility function and the upper limit constraint of the communication delay allocation, the first Lagrangian objective function can be constructed as shown in the following expression: ;(6) Simplified, we get: ;(7) in, is the Lagrange multiplier, which is the resource sensitivity of the adjustable upper bound of the total budget of end-to-end communication delay.
[0064] When constructing the first KKT condition, it is necessary to ensure that the optimization process obtains extreme values within the constraint boundary and the continuously differentiable region, and to use it to guide iterative convergence judgment.
[0065] In one embodiment, the first KKT condition includes a first-order optimality condition, a mutual relaxation condition, and a multiplier non-negativity condition.
[0066] Alternatively, the expression for the first-order optimality condition is as follows: (8) Alternatively, the mutual relaxation condition can be expressed as follows: (9) Optionally, the multiplier non-negativity condition is .
[0067] When optimizing the current communication delay allocation ratio in the external competitive game, in order to meet the system's fixed end-to-end tolerable delay constraints and achieve optimal delay resource allocation between the access network and the core network, that is, under the pre-set first KKT condition, the gradient descent method and the first Lagrangian objective function are used to solve the current communication delay allocation ratio, maximize the external joint utility function, that is, optimize the overall communication security benefits of the access network and the core network, and achieve the optimal balance of overall system performance while ensuring communication performance and security benefits, until the iteration termination condition is met and the optimized communication delay allocation ratio is obtained.
[0068] In one embodiment, when using the gradient descent method and the first Lagrangian objective function to solve the current communication delay distribution ratio, four steps are required: initializing parameters, calculating gradients and updating variables, judging convergence, and outputting the optimized communication delay distribution ratio.
[0069] Initialization parameters. Set the current communication delay allocation ratio to the initialized communication delay allocation ratio , can be initialized randomly; set the learning rate (step size) , used to control the step size of each gradient update; set the maximum number of iterations To prevent an infinite loop, set the convergence threshold , as the judgment criterion for the termination condition, that is, when the target value change of two consecutive iterations is lower than the convergence threshold, it is considered to have converged and the update is stopped.
[0070] Gradient calculation and variable update. After constructing the first Lagrangian objective function based on the external joint utility function and inequality constraints, the gradient descent method is used to iteratively optimize the current communication delay partition ratio as a variable. Assume that the expression of the current gradient is as follows: ;(10) in, It is the dynamic adjustment trend of access network and core network in the competition for delay resources.
[0071] Assume that the update rule for the k+1th iteration is expressed as follows: ;(11) in, The current communication delay allocation ratio for the kth iteration.
[0072] It is understandable that the constraints must be met after each iteration to update the current communication delay allocation ratio , and Perform projection reduction to ensure that it is always within the legal domain.
[0073] Convergence judgment. In each iterative optimization solution, when any of the iterative termination conditions is met, it is considered that the optimal solution for optimizing the communication delay partition ratio has been approached. At this time, the iteration is terminated and the optimized communication delay partition ratio is output. The iterative termination conditions include: (1) The difference between two adjacent iterations meets the convergence condition , that is, the variable update change is lower than the threshold, indicating that the system has stabilized near the optimal solution and the iteration is terminated; (2) The number of iterations reaches the maximum limit , that is, when the number of iterations exceeds the set threshold, it prevents falling into an infinite loop, stops iteration and outputs the current optimal solution; (3) Gradient close to zero condition , that is, the gradient of the external joint utility function with respect to the optimization variable tends to zero, which is considered to have reached the first-order optimality region, and the system has entered the utility equilibrium range.
[0074] Output the optimized communication delay allocation ratio. When the iteration termination conditions are met, including the current gradient or the current communication delay division ratio convergence, the communication delay allocation ratio of the current iteration is determined to be the optimal solution, that is, the optimized communication delay allocation ratio To verify whether the optimal solution has reached the Nash equilibrium state, we should further determine whether it satisfies the first-order optimality condition shown in the above expression (8), that is, , indicating that the marginal utility of both the access network and the core network is consistent under the optimized communication delay allocation ratio, and neither has the motivation to unilaterally change the strategy to make a profit. The system reaches the game equilibrium state and optimizes the communication delay allocation ratio. This is the optimal allocation solution.
[0075] To achieve the distribution ratio based on optimized communication delay The allocation scheme is deployed in the system, and the corresponding policy configuration results need to be further output: the access network finally obtains the allocable end-to-end delay budget of The communication strategy in the first current communication security strategy corresponding to the access network includes but is not limited to access point scheduling, modulation and coding level, scheduling priority, etc. The core network finally obtains an allocable end-to-end delay budget of The security policies included in the corresponding second current communication security policy include but are not limited to encryption level, authentication mechanism, access control rules and other security processing parameter sets.
[0076] The present invention provides a policy optimization method for a high-security wireless communication system based on hierarchical game. The method regards the access network and core network of the wireless communication system as independent game participants, and constructs an external joint utility function based on the external competitive game between the access network and the core network under the framework of total variable game. The method also constructs a Lagrangian objective function and a KKT condition in combination with the upper limit constraint of the communication delay allocation. With the goal of optimizing the overall communication security benefit, the method iteratively optimizes the current communication delay allocation ratio based on the gradient descent method to obtain an optimized communication delay allocation ratio. The method can achieve dynamic balanced allocation of communication delay between the access network and the core network in terms of resource share, communication delay and security protection capability, thereby ensuring the intelligent steady-state operation of the high-security wireless communication system in multiple scenarios.
[0077] Based on the above embodiment, as an optional embodiment, step S30 includes: S31, with the goal of optimizing the cooperative joint utility between the first communication decision-maker and the first security decision-maker of the access network under the optimized communication delay division ratio, iteratively optimize the first current communication security policy to obtain a first optimized communication security policy; S32. With the goal of optimizing the competitive joint utility of the second communication decision-maker and the second security decision-maker of the core network under the optimized communication delay division ratio, the second current communication security policy is iteratively solved to obtain a second optimized communication security policy.
[0078] Specifically, within the access network, the first communication decision-maker and the first security decision-maker engage in an internal cooperative game around the end-to-end optimized communication delay allocation ratio determined by the external competitive game, and use the optimized communication delay allocation ratio as a constraint condition for the optimization of the communication strategy and security strategy of the first communication decision-maker and the first security decision-maker respectively. The two parties dynamically adjust the policy configuration to achieve resource synergy and benefit balance between communication efficiency and security benefits, and output the optimal configuration results of the communication and security joint strategy under the current resource budget, providing an executable basis for dynamic scheduling and policy deployment on the access network side.
[0079] The cooperative game between the first communication decision-maker and the first security decision-maker within the access network can be modeled as a Min-Max game, in which the first communication decision-maker (Min party) aims to minimize communication delay and improve data transmission efficiency, while the first security decision-maker (Max party) attempts to enhance security protection capabilities by extending communication delay and maximize the security benefits of the system. The sum of the utility of the first communication decision-maker based on minimizing communication delay and the utility of the first security decision-maker based on maximizing security benefits is the joint cooperative utility.
[0080] Within the core network, the second communication decision-maker and the second security decision-maker compete for limited communication latency resources. The second communication decision-maker aims to minimize the additional latency introduced by security mechanisms to improve overall communication performance; the second security decision-maker focuses on enhancing security benefits. While these strategies improve protection capabilities, they typically incur higher processing latency. Nash equilibrium theory is introduced for modeling and analysis. At the Nash equilibrium point, the second communication decision-maker cannot further reduce communication latency by unilaterally weakening security policies, nor can the second security decision-maker independently enhance encryption strength without degrading overall system performance. Both parties achieve a stable balance between resource trade-offs and policy selection.
[0081] The strategic game process between the second communication decision-maker and the second security decision-maker within the core network forms an optimization structure with opposing goals and mutual constraints under the constraint of a limited communication delay budget, resulting in an antagonistic relationship between communication performance and security benefits, with one increasing while the other decreases.
[0082] Therefore, the competitive game process involving the strategic interactions between the second communication decision-maker and the second security decision-maker within the access network can be modeled as a zero-sum game. The utility functions of the second communication decision-maker and the second security decision-maker are inversely proportional. An increase in the strategic benefits of either party inevitably results in a loss for the other. The overall utility of the system is conserved, driving system performance toward dynamic equilibrium during the confrontation. When the sum of the utilities of the second communication decision-maker and the second security decision-maker is zero, the competitive joint utility is optimal.
[0083] When the optimized communication delay division ratio is obtained in the external competition game, the internal cooperation game of the access network and the internal competition game of the core network are respectively carried out based on the optimized communication delay division ratio.
[0084] In the internal cooperative game of the access network, that is, in step S31, the first current communication security policy of the access network is iteratively optimized with the goal of optimizing the cooperative joint utility under the optimized communication delay division ratio obtained by the first communication decision-maker and the first security decision-maker in the external competitive game to obtain a first optimized communication security policy.
[0085] In the internal competition game of the core network, that is, in step S32, the second current communication security strategy of the core network is iteratively solved with the goal of optimizing the competitive joint utility of the second communication decision-maker and the second security decision-maker under the optimized communication delay division ratio obtained in the external competition game to obtain the second optimized communication security strategy.
[0086] Optionally, the execution order of step S31 and step S32 may be simultaneous execution, step S31 may be executed first, and then step S32, or step S32 may be executed first, and then step S31.
[0087] The optimized communication delay allocation ratio determined by the external competitive game provides clear resource budget boundaries for the access and core networks. The access and core networks then independently optimize communication and security policies within their respective available resources. Although the access and core networks optimize communication security policies independently, they collaborate through a unified delay budget framework to achieve a secure communication balance across the entire system.
[0088] The present invention provides a policy optimization method for a high-security wireless communication system based on hierarchical game. After obtaining the optimized communication delay division ratio through external competitive game, a cooperative game model is constructed according to the characteristics of the access network. The communication security policy of the access network is iteratively optimized with the goal of optimizing the cooperative joint utility of the communication decision-makers and security decision-makers of the access network under the optimized communication delay division ratio. A competitive game model is constructed according to the characteristics of the core network. The communication security policy of the core network is iteratively optimized with the goal of optimizing the competitive joint utility of the communication decision-makers and security decision-makers of the core network under the optimized communication delay division ratio. This helps establish a continuously iterative global optimization interaction mechanism to solve the optimal dynamic allocation strategy of delay resources, and solves the optimal global communication security policy of the wireless communication system, so as to achieve dynamic game equilibrium between the access network and the core network, thereby achieving the overall synergy between the communication efficiency and security assurance capability of the wireless communication system.
[0089] Based on the above embodiment, as an optional embodiment, the first current communication security policy includes the communication policies of the first communication decision-maker and the security policies of the first security decision-maker; the cooperative joint utility includes the communication decision-maker utility and the security decision-maker utility; step S31 includes: Under a second KKT condition, with the goal of optimizing the cooperative joint utility of the access network, the first current communication security policy is iteratively optimized based on a policy update rule and a second Lagrangian objective function to obtain the first optimized communication security policy; the policy update rule includes sub-update rules for each communication policy and each security policy in the first current communication security policy; The second Lagrangian objective function includes a cooperative joint utility function, an access network communication delay budget constraint, a bandwidth resource constraint, and a security constraint; The cooperative joint utility function is determined based on the communication decision-maker utility function and the security decision-maker utility function; the communication decision-maker utility function is determined based on the policy benefits of each communication policy of the first communication decision-maker and the delay overhead of each security policy of the first security decision-maker; the security decision-maker utility function is determined based on the anti-attack capability of each security policy of the first security decision-maker and the security synergy benefit of each communication policy of the first communication decision-maker; The access network communication delay budget constraint is determined based on the optimized communication delay division ratio and the access network delay budget sensitivity; the bandwidth resource constraint is determined based on the access network allocated bandwidth, the total system available bandwidth and the bandwidth resource sensitivity in the first current communication strategy; the security constraint is determined based on the minimum access network security utility, the security decision-maker utility and the security sensitivity; the Lagrangian multiplier of the second Lagrangian objective function includes the access network delay budget sensitivity, the bandwidth resource sensitivity and the security sensitivity.
[0090] The communication strategies of the first communication decision-maker include but are not limited to allocating bandwidth to the access network. , the transmitting power of the communication party and access point selection strategy At least one of the above can constitute the first communication strategy set of the first communication decision-maker .
[0091] The security policies of the first security decision-maker include but are not limited to wireless data flow encryption policies , integrity protection strategy , device trusted authentication strategy and Privacy Policy At least one of the above can constitute the first security policy set of the first security decision-maker .
[0092] The communication strategies of the first communication decision-maker and the security strategies of the first security decision-maker are used as decision variables, respectively representing the strategy space of the first communication decision-maker and the first security decision-maker in the internal cooperative game, laying the foundation for subsequent strategy collaborative optimization modeling and solution.
[0093] Specifically, before the internal cooperative game between the first communication decision-maker and the first security decision-maker begins, a Lagrangian objective function for the access network's internal cooperative game, also known as the second Lagrangian objective function, must be pre-established. Furthermore, to ensure that the resulting policy solution under these constraints is a local optimal solution, the internal cooperative game incorporates the KKT optimality condition, also known as the second KKT condition.
[0094] When constructing the second Lagrangian objective function, in order to quantitatively characterize the optimization objectives of the first communication decision-maker and the first security decision-maker of the access network under the resource constraint of optimizing the communication delay allocation ratio, the communication decision-maker utility function of the first communication decision-maker is first constructed. and the security decision-maker utility function of the first security decision-maker , and based on this, define the cooperative joint utility function within the access network , to measure the utility of communication decision-makers.
[0095] In one embodiment, the utility function of the communication decision-maker is expressed as follows: ;(12) in, The utility of the communication decision-maker is used to measure the communication strategy and security policies The average delay overhead generated by the communication process; is the size of the service data packet; Selecting a policy for an access point The channel gain brought by is the noise power; The delay overhead of each security policy of the first security decision-maker; Additional overhead caused by access point switching.
[0096] It can be understood that the policy benefits of each communication policy of the first communication decision-maker are based on the access point selection policy. Channel gain , additional overhead caused by access point switching , access network bandwidth allocation and the transmitting power of the communication party jointly constituted.
[0097] The optimization goal of the first communication decision-maker is to minimize the communication delay to improve the data transmission efficiency, that is, to minimize .
[0098] In one embodiment, the expression of the security decision-maker utility function is as follows: ; (13) ;(14) ;(15) in, It is the security decision-maker utility, used to measure the communication security under the current security policy configuration; The number of security policy items of the first security decision-maker; is the attack success probability of the k-th security policy; is the anti-attack capability of the k-th security policy; The security resource investment level for the kth security policy, such as encryption strength, authentication level, etc. The security synergy benefits of the communication strategy of the first communication decision-maker; is the channel signal-to-noise ratio (SNR); Normalize the bandwidth ratio; It is the security factor assigned to different access points.
[0099] Among them, the probability of attack success The design of this paper is to incorporate the indirect impact of each access network communication strategy on the attack risk into the modeling of the security decision-maker's utility.
[0100] The optimization goal of the first security decision-maker is to maximize the communication delay and security benefits to achieve the optimal balance between transmission reliability and security, that is, to minimize .
[0101] In one embodiment, the expression of the cooperative joint utility function is as follows: ; (16) in, is the weight coefficient of the communication decision-maker's utility; is the weight coefficient of the security decision-maker's utility.
[0102] Among them, the weight coefficients of the communication decision-maker utility and the security decision-maker utility are used to control the importance ratio of the first communication decision-maker and the first security decision-maker in the cooperative joint optimization. The weighted cooperative joint utility function constructed in this way can achieve a trade-off balance between access network communication efficiency and security performance.
[0103] The optimization goal of the access network is to achieve a balance between communication efficiency and security performance. The optimization goal of the cooperative joint utility constructed is .
[0104] When constructing the second Lagrangian objective function, three constraints must be introduced to ensure the feasibility of the optimization problem and meet the system's communication quality of service (QoS) requirements: (1) Access network communication delay budget constraint , ensuring that the communication delay of the access network does not exceed the communication delay budget determined by the optimized communication delay allocation ratio determined based on the external competition game; (2) Bandwidth resource constraints , ensuring bandwidth allocation in the access network The total available bandwidth of the wireless communication system shall not be exceeded ; (3) Security constraints , ensuring the effectiveness of access network security decision-making Not less than the minimum access network security utility , ensuring the strength of safety protection.
[0105] Based on the cooperative joint utility function, access network communication delay budget constraint, bandwidth resource constraint, and security constraint, the expression of the second Lagrangian objective function is as follows: ; (17) in, 、 、 are all Lagrange multipliers; Budget sensitivity for access network delay; is the bandwidth resource sensitivity; For security sensitivity.
[0106] On the basis of the second Lagrangian objective function, with the goal of optimizing the cooperative joint utility of the access network, the first current communication security strategy of the first communication decision-maker and the first security decision-maker is jointly solved to make the second Lagrangian objective function The optimal solution is achieved under various constraints, thereby achieving the optimal balanced configuration of communication efficiency and security effectiveness on the access network side.
[0107] In one embodiment, the second KKT condition includes a first-order optimality condition, a mutual relaxation condition, and a multiplier non-negativity condition.
[0108] Optionally, the first-order optimality condition is used to ensure that the gradient direction of the Lagrangian objective function in the feasible region is zero, thereby achieving the necessary conditions for the optimal solution. Its expression is as follows: (18) Optionally, the mutual relaxation condition is used to indicate that when a constraint is not compactly activated, its corresponding Lagrange multiplier is zero, reflecting the coupling relationship between the constraint and the multiplier. Its expressions are shown in the following equations (19)-(21): ; (19) ; (20) .(twenty one) Optionally, the multiplier non-negativity condition is 、 、 , which is used to ensure that all multipliers meet the non-negativity requirement, thereby maintaining the physical meaning of the Lagrange multipliers and the constraint duality conditions.
[0109] In the cooperation game within the access network, the first communication decision-maker and the first security decision-maker focus on the first current communication security strategy. An adversarial optimization is carried out with the goal of optimizing the cooperative joint utility. The cooperative joint utility function in the game between the first communication decision-maker and the first security decision-maker is iteratively solved through the minimax optimization mechanism until the iterative termination condition is met, and converges to an optimal joint strategy solution that is dynamically balanced between communication efficiency and security benefits, thus obtaining the first optimized communication security strategy.
[0110] In one embodiment, when iteratively optimizing the first current communication security policy based on the policy update rule and the second Lagrangian objective function, four steps are required: initializing parameters, calculating gradients and updating variables, convergence judgment, and outputting the first optimized communication security policy.
[0111] Initialization parameters. Set the communication strategy set that can be adjusted by the first communication decision-maker to , including initial bandwidth allocation , transmit power , access point selection , reflecting the initial state of communication performance optimization. Set the initial security policy set that can be adjusted by the first security decision-maker , including initial encryption strength , integrity protection mechanism , trusted authentication mechanism and privacy protection measures , reflecting the combination of activation strategies for the initial security configuration of the system. Setting the Lagrange multiplier . Set the initial learning rate , to control the step size of variable update in each iteration. Set the tolerance threshold , a precision control parameter to determine whether the optimization has converged. Set the maximum number of iterations .
[0112] Calculate the gradient and update the variables. In the kth iteration, iteratively optimize the first current communication security policy based on the policy update rule and the second Lagrangian objective function. For different communication policies, security policies, and Lagrangian multiplier types, a joint optimization update is performed according to a predetermined sub-update rule.
[0113] Communication strategy update. During each round of internal game optimization, the communication decision-maker updates the communication strategy variables according to the second Lagrangian objective function to minimize the system loss caused by communication delay. The sub-update rule for the communication strategy variables, which are continuous variables, is based on the gradient descent method. The update rule for the k+1th iteration is expressed as follows: .(twenty two) For example, for the continuous variables of access network allocated bandwidth and communication party transmit power, their update formulas are shown in the following equations (23) and (24), respectively: ;(twenty three) ;(twenty four) in, is the learning rate, which is used to control the strategy update step size to ensure that the communication delay is gradually reduced and the convergence is stable.
[0114] The update of the access point selection strategy as a discrete variable cannot use gradient descent. Its sub-update rule is updated based on greedy search or minimum delay criterion. The expression of the update rule for the k+1th iteration is as follows: (25) The sub-update rule of the security policy variable as a continuous variable is guided by the goal of maximizing the system security. The gradient direction makes the security benefit continuously increase. The update is based on the gradient ascent method. The expression of the update rule for the k+1th iteration is as follows: (26) For example, for the wireless data stream encryption strategy, integrity protection strategy, device trusted authentication strategy, and privacy protection strategy of continuous variables, their update formulas are shown in the following equations (27), (28), (29), and (30), respectively: ; (27) ; (28) ;(29) (30) Lagrange multiplier update. In each iteration, to ensure that the access network communication strategy meets the system constraints, the access network delay budget sensitivity is introduced. , bandwidth resource sensitivity and security sensitivity The iterative update methods are shown in the following equations (31)(32)(33): ; (31) ; (32) ; (33) in, is the total bandwidth of the communication strategy in the kth iteration process; is the access network security benefit calculated in the k-th iteration process; The default minimum safety performance requirements; It is a non-negative truncation operation, which is used to force the multiplier to always satisfy the non-negativity.
[0115] By adaptively adjusting the above-mentioned Lagrange multipliers, it is possible to optimize communication and security strategies while ensuring that system operation always complies with resource and security constraints, thereby improving the controllability and robustness of the overall operation.
[0116] Convergence determination. After completing each round of access network policy variable and Lagrange multiplier updates, the system determines whether the optimal solution has been reached or the iteration termination conditions have been met. When any of the iteration termination conditions are met, the iteration process is terminated and the first optimized communication security policy is output. The iteration termination conditions include: (1) The change in the communication strategy variable of the first communication decision-maker is less than the preset tolerance threshold ; (2) Reaching the set maximum number of iterations .
[0117] Output the first optimized communication security strategy. When the iteration termination conditions are met, including the convergence of the communication security strategy, the communication security strategy of the current iteration is determined to be the optimal solution of the cooperation game within the access network, which is the first optimized communication security strategy. , that is Optimized communication strategy for the first communication decision-maker, Optimized security strategy for the first security decision maker. First optimized communication security strategy , which satisfies the optimal balance between minimizing communication delay and maximizing security benefits, and at the same time complies with the KKT optimality conditions under Lagrangian constraints.
[0118] The present invention provides a policy optimization method for a high-security wireless communication system based on hierarchical game. The method regards the communication party and the security party of the access network as independent game participants, and constructs a cooperative joint utility function based on the internal cooperative game of the access network under the Min-Max game framework, with the cooperative joint utility jointly determined by the communication decision utility of the communication party and the security decision utility of the security party. The Lagrangian objective function and the KKT condition are constructed in combination with the access network communication delay budget constraint, bandwidth resource constraint and security constraint. With the optimization of the cooperative joint utility as the goal, the first current communication security policy is iteratively optimized based on the update rules set based on the characteristics of each communication security policy to obtain a first optimized communication security policy. The method can realize dynamic balanced allocation of the access network under a given communication delay budget, and ensure the intelligent steady-state operation of the high-security wireless communication system in multiple scenarios.
[0119] Based on the above embodiment, as an optional embodiment, the step of iteratively optimizing the first current communication security policy based on a policy update rule and a second Lagrangian objective function with the goal of optimizing the cooperative joint utility of the access network to obtain the first optimized communication security policy includes: constructing a first input vector based on the first current communication security policy; Inputting the first input vector into a first policy update model to obtain a first optimized communication security policy output by the first policy update model; the first policy update model is implemented based on a deep Q network; The state space of the first policy update model is determined based on the communication strategies and security strategies in the first current communication security strategy; the action space of the first policy update model is determined based on the policy update rule; and the reward function of the first policy update model is determined based on the cooperative joint utility function in the second Lagrangian objective function.
[0120] Considering that the cooperative game within the access network is a nonlinear and complex optimization problem involving discrete variables (such as security policy combinations and access point selection) and continuous variables (such as bandwidth allocation and transmit power), and that these policy variables are highly coupled, traditional analytical methods are difficult to solve efficiently, or even infeasible. To solve this joint policy optimization problem, a mathematical optimization method is combined with a heuristic optimization algorithm to dynamically jointly optimize the policy configurations of the first communication decision-maker and the first security decision-maker, ultimately converging on a joint policy combination that achieves the optimal trade-off between communication performance and security benefits.
[0121] Specifically, a first policy update model is constructed in advance based on a deep Q-network (DQN), the state space of the first policy update model is determined according to each communication strategy and each security strategy in the first current communication security strategy, the action space of the first policy update model is determined according to the policy update rules of each communication strategy and each security strategy, and the reward function of the first current communication security strategy is determined according to the cooperative joint utility function in the second Lagrangian objective function constructed above.
[0122] Specifically, the state space of the first strategy update model is set , indicating the current communication resource status, security policy configuration and environmental situation information of the access network (such as bandwidth, power, access point channel quality, historical feedback, etc.).
[0123] Set the action space of the first policy update model , which represents the fine-tuning operation of each round of strategy (increase or decrease, replacement, etc.).
[0124] By using a cooperative joint utility function that considers both communication performance and security benefits as a reward function, intelligent joint scheduling of communication security strategies can be achieved. The expression of the reward function is as follows: ; (34) in, In state Take strategic actions Immediate rewards after is the communication decision-maker utility of the first communication decision-maker. The smaller the better, so a negative weight is introduced. The security decision-maker utility of the first security decision-maker is the larger the better; 、 Utility for communication decision-makers and security decision-making effectiveness The weight coefficient satisfies the relative importance configuration in the system scenario.
[0125] After building the first policy update model, initialize the network structure and parameters of the first policy update model. As the first policy update model built based on the deep Q network, including the main Q network and the target Q network, the model parameters of the main Q network and the target Q network are initially set to be consistent, and then synchronized and updated every certain number of steps. Set the training parameters of the first policy update model, including: the discount factor used to balance current benefits and future returns ; The learning rate used to control the gradient update step size ; The initial value of the exploration rate is set to 1, and it decays linearly or exponentially to the minimum value as the training progresses; the experience replay pool used to store state transition samples The capacity is set to 10,000 to 50,000; the batch size used to sample data from the experience pool to update the network in each round of training is set to 64 or 128.
[0126] After pre-constructing and completing the pre-training of the first policy update model, in the process of cooperative game within the access network, the first current communication security policy is used to construct a first input vector, and the first input vector is input into the pre-trained first policy update model to obtain the first optimized communication security policy output by the first policy update model. .
[0127] The first optimized communication security strategy includes a first optimized communication strategy and the first optimized security strategy , it is judged that the Nash equilibrium state is reached, that is, the first-order optimality condition is met, as well as , which means that neither the first communication decision-maker nor the first security decision-maker can improve their own benefits by unilaterally changing the strategy, and the system communication delay and security benefits reach a stable balance state. It is considered that the Nash equilibrium solution is achieved, and the first optimized communication security strategy This is the optimal secure communication configuration result obtained through the internal cooperative game of the access network under reinforcement learning optimization.
[0128] The present invention provides a policy optimization method for a high-security wireless communication system based on hierarchical game. By combining mathematical optimization methods and heuristic optimization algorithms, a policy update model is constructed based on a deep Q network. Communication performance and security performance are taken into consideration, and the reward function of the policy update model is determined according to the cooperative joint utility function in the second Lagrangian objective function. In the nonlinear and complex optimization problem of cooperative game within the access network, the game feedback and environmental state are combined to gradually learn the optimal policy combination, and the intelligent dynamic joint optimization is solved to obtain an optimized communication security policy that converges to achieve the optimal trade-off between communication performance and security benefits.
[0129] In one embodiment, pre-training of the first policy update model is completed through the steps of interactive sampling and experience replay, environment interaction and sample recording, network training and parameter update, training termination judgment and policy convergence judgment: Interactive sampling and experience replay. At each reinforcement learning training time step, a policy action is selected based on the current system state, and corresponding environmental interactions are performed. Specifically, this involves adopting a greedy strategy to balance exploration and exploitation. First, the policy updates the model to select an action from the set of possible actions using the current probability (i.e., the exploration rate). Simultaneously, the action with the highest valuation in the current Q network is selected with a probability of "1-exploration rate." This involves calculating the predicted utility value of each candidate action in the current system state using the policy function, thereby determining the optimal behavior.
[0130] Environment interaction and sample recording. After executing the selected optimal action, the state of the first policy update model transitions to the next state, and an immediate reward is obtained based on the reward function calculated based on the cooperative joint utility function. The experience quadruple formed by the state transition process is stored in the experience replay pool for subsequent sample sampling and valuation correction in Q network training.
[0131] Network training and parameter update. After the experience replay pool is filled to the preset capacity, the first policy update model enters the policy network training phase. Through small batch sampling and gradient optimization mechanism, the Q network parameters are continuously updated to approach the optimal policy function. First, a small batch of four-tuple samples are randomly sampled from the experience replay pool for the current training round. Then, the loss function in the form of Bellman error is used as the objective function to adjust the parameters of the main Q network. Optimize and get the target value , the predicted value is The first strategy update model uses stochastic gradient descent to minimize the above loss function and complete the update of the main Q network parameters. Then, after a certain number of training steps (such as 1000 rounds), the main network parameters are synchronized to the target network, that is, The aforementioned parameter synchronization strategy is used to enhance the stability of the training process, avoid Q-value update oscillations, and help strengthen the convergence and speed of the learning process. Through network training and parameter update steps, the first policy update model can gradually approach the optimal policy through continuous iteration, effectively outputting a safety function combination solution under joint optimization.
[0132] Specifically, the expression of the objective function is as follows: ; (35) in, Because the reward function Calculated instant rewards; is the discount factor, which is used to measure the importance of future returns; are the parameters of the current main Q network; are the parameters of the target network.
[0133] Training termination judgment and strategy convergence judgment. In the first strategy update model training process, in order to improve the stability and accuracy of strategy decision-making, the following two types of termination conditions are introduced: First, the exploration rate of the greedy strategy is dynamically updated and gradually reduced in a linear decreasing manner to reduce the proportion of random exploration and enhance the certainty of the model strategy output; second, the training convergence judgment condition is set. When the difference in the Q network output in the two rounds of training meets the convergence threshold, , or when the number of training rounds reaches the set maximum number of training rounds (such as 1000 rounds), the training is terminated. At the same time, by real-time monitoring of key performance indicators such as communication decision-making party utility and security decision-making effectiveness , ensuring that the performance of the first strategy update model converges stably to the optimal state during training.
[0134] Based on the above embodiment, as an optional embodiment, when the second current communication security policy is a discrete policy, step S32 includes: Based on the second current communication security policy, defining a communication hybrid policy set of a second communication decision-maker of the core network and a security hybrid policy set of a second security decision-maker of the core network; defining a first hybrid strategy probability distribution variable corresponding to the communication hybrid strategy set and a second hybrid strategy probability distribution variable corresponding to the security hybrid strategy set; Constructing a game matrix; each element in the game matrix is a loss function value obtained by substituting a communication hybrid strategy from the communication hybrid strategy set and a security hybrid strategy from the security hybrid strategy set into a pre-constructed communication party loss function; the communication party loss function is determined based on the core network end-to-end delay and the security policy communication overhead; Based on the optimized communication delay partition ratio, the game matrix, and the first mixed strategy probability distribution variable, with the goal of minimizing the upper limit of communication expected loss, constructing the original linear programming model of the second communication decision-maker; Based on the optimized communication delay partition ratio, the game matrix, and the second mixed strategy probability distribution variable, a dual linear programming model of the second security decision-maker is constructed with the goal of maximizing the lower bound of the expected security benefit; Calling a standard linear programming solver to jointly solve the original linear programming model and the dual linear programming model to obtain an optimal solution pair of the hybrid strategy probability distribution with the optimal competitive joint utility; The second optimized communication security strategy is determined based on the optimal solution pair of the hybrid strategy probability distribution.
[0135] Considering the diversity and feasibility of the core network policy space, two types of solution methods can be designed based on the different types of policy spaces. In the case where the second current communication security policy is a discrete policy, the policy space is a discrete policy set. In this discrete policy space, policies are given in enumerated form, indicating that each party chooses the optimal combination from a limited set of options. A linear programming (LP) approach is used to construct a minimax adversarial optimization model to directly solve for the probability distribution of the optimal hybrid strategy of the second communication decision-maker and the second security decision-maker.
[0136] The communication strategies for the second communication decision-maker include but are not limited to allocating bandwidth to the access network. , the transmitting power of the communication party and access point selection strategy At least one of the above can constitute the first communication strategy set of the first communication decision-maker .
[0137] The security policies of the first security decision-maker include but are not limited to wireless data flow encryption policies , integrity protection strategy , device trusted authentication strategy and Privacy Policy At least one of the above can constitute the first security policy set of the first security decision-maker .
[0138] Specifically, in the second current communication security policy In the case of discrete strategies, first build a communication hybrid strategy set that defines the second communication decision-maker of the core network A security hybrid policy set with the second security decision-maker of the core network , where each 、 They represent a set of specific communication hybrid strategies and security hybrid strategies, respectively. It should be noted that the hybrid strategy means that the second communication decision-maker and the second security decision-maker, as game players, do not choose a single strategy, but instead mix and play strategies in their strategy sets according to certain probability weights.
[0139] Further define the hybrid strategy probability distribution variable and satisfy the probability normalization and non-negativity constraints, including defining the first hybrid strategy probability distribution variable corresponding to the communication hybrid strategy set , indicating that the second communication decision-maker uses the probability Adopt a communication strategy ; and define the second hybrid strategy probability distribution variable corresponding to the safe hybrid strategy set , indicating that the second security decision-maker uses probability Adopt a security strategy .
[0140] Before further constructing the game matrix, it is necessary to first determine the communication party loss function determined by the second communication decision-maker in the core network based on the core network end-to-end delay and security policy communication overhead.
[0141] The second communication decision-maker aims to minimize the communication performance loss. Under this condition, the expression of the communication party loss function of the second communication decision-maker is as follows: ; (36) in, is the loss function value; The core network end-to-end delay caused by the combination of communication security policies; Security policy communication overhead refers to the additional overhead on communication resources (such as encryption and authentication) introduced by security policies.
[0142] It can be understood that within the core network, the communication party loss function is a negative value, which means that the smaller the loss, the higher the communication efficiency.
[0143] At the same time, the second security decision-maker aims to maximize security benefits in the strategy combination Under this condition, the expression of the security party profit function of the second security decision-maker is as follows: ; (37) in, is the profit function value; For security policy benefits; Encryption processing delays for security policies; The weight coefficient for balancing security benefits and latency impact.
[0144] It can be understood that within the core network, the security benefit function is positive, indicating that the higher the benefit, the better the system security.
[0145] It should be noted that the second communication decision-maker and the second security decision-maker within the core network are in an internal competitive game. The second communication decision-maker and the second security decision-maker engage in strategic confrontation around the target difference between communication performance and security benefits. This represents that the internal competitive game of the core network belongs to a strictly opposing zero-sum game structure. The loss of the communication party is the opposite of the gain of the security party, that is, it satisfies The symmetrical relationship between the two is that the smaller the loss of the second communication decision-maker, the lower the benefit of the second security decision-maker, and vice versa, which reflects the conservation of the total utility of the game system.
[0146] Based on the defined communication party loss function and security party benefit function, we further construct a game matrix for linear programming solution that describes the utility confrontation relationship between the second communication decision-maker and the second security decision-maker under discrete strategy combinations. ,and . Among them, the row index Represents a communication hybrid strategy , column index Represents a secure hybrid policy , the first The element indicates that the second communication decision-maker has The loss function value under , that is, each element in the game matrix is the loss function value obtained by substituting a communication mixed strategy in the communication mixed strategy set and a security mixed strategy in the security mixed strategy set into the communication party's loss function.
[0147] Understandably, since the competition within the core network is a zero-sum game structure, The symmetry identity relation of .
[0148] Further construct the original linear programming model of the second communication decision-maker. Under the premise of facing the most unfavorable response of the second security decision-maker, the second communication decision-maker aims to minimize the upper limit of the expected loss of communication. At this time, based on the optimized communication delay partition ratio obtained by the external competitive game, the constructed game matrix and the first mixed strategy probability distribution variable, the original linear programming model shown in the following expression is constructed: ; (38) in, Mixed strategies for communication The maximum communication delay that may occur under the most unfavorable security hybrid strategy; is the number of communication hybrid strategies in the communication hybrid strategy set.
[0149] By constructing the original linear programming model, it can ensure that the second communication decision-maker can face any security mixed strategy. In the case of , thereby minimizing the worst case scenario.
[0150] Further construct the dual linear programming model of the second security decision-maker. Under the premise of facing the most unfavorable response of the second communication decision-maker, the second security decision-maker aims to maximize the lower limit of the expected security benefit. At this time, based on the optimized communication delay partition ratio obtained from the external competitive game, the constructed game matrix and the second mixed strategy probability distribution variable, a dual linear programming model is constructed as shown in the following expression: ; (39) in, Hybrid strategy for security The encryption processing delay or system processing load introduced; is the number of security mix policies in the security mix policy set.
[0151] By constructing a dual linear programming model, it is possible to ensure that the second security decision-maker can face any communication hybrid strategy. In the case of , thereby minimizing the worst case scenario.
[0152] After constructing the original linear programming model and the dual linear programming model, the standard linear programming solver is called to jointly solve the original linear programming model and the dual linear programming model to obtain the optimal solution of the hybrid strategy probability distribution with the optimal competitive joint utility that satisfies the maximum-minimum symmetric structure and characterizes the internal competition game of the core network. , and based on the optimal solution of the hybrid strategy probability distribution, perform joint sampling to generate and determine the second optimized communication security strategy .
[0153] It can be understood that the optimization goal of the second communication decision-maker is to minimize the upper limit of the communication expected loss (maximum loss), so the optimized communication hybrid strategy in the second optimized communication security strategy can be defined as , that is, under the premise of uncertainty of the security hybrid strategy, the second communication decision-maker selects the optimal strategy that minimizes its own loss upper limit from its strategy set; the optimization goal of the second security decision-maker is to maximize the lower limit of the security expected benefit, so the optimized security hybrid strategy in the second optimized communication security strategy can be defined as , that is, under the premise of uncertainty of the communication hybrid strategy, the second security decision-maker selects the optimal strategy that maximizes the lower limit of its own benefits from its strategy set.
[0154] It should be noted that the optimal solution of the hybrid strategy probability distribution obtained by jointly solving the original linear programming model and the dual linear programming model is , the game equilibrium value can be calculated , the game equilibrium value is used to indicate that the communication party and the security party form a stable utility level in the confrontation. In this state, the system cannot obtain a better result through unilateral adjustment, which reflects the stability and optimality of the game. More specifically, the game equilibrium value It represents the expected utility equilibrium level of the communication decision-maker and the security decision-maker in the core network under Nash equilibrium, that is, when the second communication decision-maker adopts the strategy When the maximum expected loss of the second security decision-maker under any combination of response strategies does not exceed the equilibrium utility value ; Adopt strategies at the second security decision-making party When the minimum expected benefit of the second communication decision-maker under any combination of response strategies is not less than the equilibrium utility value .
[0155] It is understandable that in order to enhance the engineering feasibility of the strategy solution, certain constraints must be followed during the model solution process, that is, the overall communication delay overhead of the second optimized communication security strategy obtained by the solution must be less than the communication delay budget determined by the optimized communication delay division ratio based on the output of the external competitive game. , that is .
[0156] The present invention provides a strategy optimization method for a high-security wireless communication system based on hierarchical games. When faced with an internal competitive game of a core network based on a zero-sum game framework, the method models the game as a minimax problem, and adopts a linear programming method for modeling and solving the problem. The optimization goal is to minimize the upper limit of the communication expected loss of the communication decision-maker and maximize the lower limit of the security expected benefit of the security decision-maker. The original linear programming model and the dual linear programming model are constructed, and the optimal strategy combination of the two game players under the most unfavorable circumstances is jointly solved. The optimal solution pair of the mixed strategy probability distribution that achieves Nash equilibrium and optimal competitive joint utility is obtained, and a strategy solution scheme that unifies communication delay, resource overhead and security benefit is also realized.
[0157] Based on the above embodiment, as an optional embodiment, when the second current communication security policy is a continuous policy, step S32 includes: constructing a second input vector based on the communication policy in the second current communication security policy; Inputting the second input vector into a second policy update model to obtain an optimized communication policy output by the second policy update model; the second policy update model is implemented based on a deep Q network; constructing a third input vector based on the security policy in the second current communication security policy; Inputting the third input vector into a third policy update model to obtain an optimized security policy output by the third policy update model; the third policy update model is implemented based on a deep Q network; Based on the optimized communication strategy and the optimized security strategy, obtaining the second optimized communication security strategy; The state spaces of the second policy update model and the third policy update model are both determined based on each communication policy and each security policy in the second current communication security policy; The action space of the second policy update model is determined based on the communication policy in the second current communication security policy; the reward function of the second policy update model is determined based on the core network end-to-end delay and the security policy communication overhead; The action space of the third policy update model is determined based on the security policy in the second current communication security policy; the reward function of the third policy update model is determined by the security policy benefit and the security policy encryption processing delay.
[0158] In the second case where the current communication security strategy is a continuous strategy, the strategy space is a continuous strategy set. The strategy variables in the continuous strategy space are continuous real numbers, such as bandwidth share, encryption coefficient, delay threshold, etc. Since the traditional LP method is only applicable to discrete variables, it is difficult to directly model and solve the minimax optimization problem under the continuous strategy. At this time, the DQN method based on deep reinforcement learning is adopted to iteratively approximate the minimax equilibrium strategy through the reinforcement training process to deal with complex strategy optimization problems such as nonlinearity, high dimension, and non-differentiable.
[0159] Specifically, a second policy update model for optimizing the communication policy of the second communication decision-maker and a third policy update model for optimizing the security policy of the second security decision-maker are constructed in advance based on the deep Q network, and the state space of the second policy update model and the third policy update model are jointly determined according to each communication policy and each security policy in the second current communication security policy. , which can describe the current core network communication environment and security situation. The joint state at each moment is expressed as .
[0160] Determine the action space of the second policy update model based on the communication policy in the second current communication security policy , based on the security policy in the second current communication security policy, determine the action space of the third policy update model , the strategic action of the second communication decision-maker In the continuous interval Internal value, policy action of the second security decision-maker It also takes values in a continuous interval.
[0161] Based on the states and actions of the second policy update model and the third policy update model, a state transition function between the second communication decision maker and the second security decision maker for characterizing the joint state of the core network can be determined. Specifically, the second communication decision maker and the second security decision maker are respectively based on the current state , using their respective strategy functions 、 Select the corresponding action 、 , working together with the environment to generate the next state And the corresponding reward value. Strategy function 、 That is, the action decision rule of the second communication decision-maker and the second security decision-maker in the current joint state. It is guided by the action value function output by the current Q network of the second policy update model and the third policy update model, and combined with - Greedy mechanism determines action selection: based on probability Select the action with the largest current Q value, with probability Randomly explore other possible actions to maintain exploratory nature in the early stages of training and enhance exploitation during the convergence phase.
[0162] At this time, the expressions of the state transition functions of the second strategy update model and the third strategy update model are as follows: ; (40) in, is the state transfer function, used to characterize the current joint state The transfer result under the joint strategy of the first communication decision-maker and the first security decision-maker.
[0163] Furthermore, the reward functions of the second policy update model and the third policy update model are respectively constructed based on the profit and loss function. The reward function of the second policy update model is used to guide the second communication decision-maker to minimize the performance, and the reward function of the third policy update model is used to guide the second security decision-maker to maximize the security benefit.
[0164] Optionally, the reward function of the second policy update model is determined based on the core network end-to-end delay and the security policy communication overhead. The expression of the reward function of the second policy update model is as follows: ;(41) in, The end-to-end latency of the core network; Security policy communication overhead.
[0165] Optionally, the reward function of the third policy update model is determined by the security policy benefit and the security policy encryption processing delay. The expression of the reward function of the third policy update model is as follows: ; (42) in, For security policy benefits; Encryption processing delay for security policy.
[0166] Further initialize the Q network parameters and structure to estimate the joint state , the expected return value function after taking each strategy action.
[0167] The neural network parameter vectors of the second and third strategy update models are 、 , the target Q network parameters corresponding to the Xavier initialization method are uniformly adopted as 、 ; Set the training parameters including the discount factor to , used to balance current gains and future returns; learning rate , used to control the gradient update step size; exploration rate The initial value is set to 1 and decays linearly or exponentially to the minimum value as the training progresses; the experience replay pool capacity Set to 10000~50000 to store state transition samples; batch size is set to 64 or 128 to sample data from the experience pool in each round of training to update the network.
[0168] When specifically optimizing the second current communication security policy of the continuous policy, the second input vector is constructed using the communication policy in the second current communication security policy, and the third input vector is constructed using the security policy in the second current communication security policy. The second input vector is input into the second policy update model, and the third input vector is input into the third policy update model to obtain the optimized communication policy and the optimized security policy.
[0169] In one embodiment, the second policy update model and the third policy update model are trained based on the following method: The Q network of the second policy update model is expressed as , the Q network of the third strategy update model is expressed as ; Accordingly, the target value of the second communication decision-maker is expressed as , the target value of the second security decision maker is expressed as .
[0170] Optionally, the Q network is constructed using a Multi-Layer Perceptron (MLP) architecture, which includes an input layer, several hidden layers (such as ReLU activation functions), and an output layer that returns the Q value prediction.
[0171] The training process of the second strategy update model and the third strategy update model includes interactive sampling and experience playback steps, environment interaction and sample recording steps, network training and parameter updating steps, and training termination judgment and strategy convergence judgment steps.
[0172] In the interactive sampling and experience replay steps, the second strategy update model and the third strategy update model are in the system state Under the current Q network, the evaluation value is output and the action is selected using the ε-greedy strategy. 、 , combined to form the current strategy combination The system environment executes the corresponding deployment behavior based on the action and returns the state transfer result. , and the instant reward value corresponding to the communication party and the security party 、 , the reward function is defined by the aforementioned communication performance loss and security benefit function.
[0173] In the environmental interaction and sample recording steps, an interaction process is organized into a six-tuple sample in the following form: , and stored in the experience replay pools of the two models. The experience replay mechanism improves data utilization efficiency, enhances training stability, and effectively mitigates correlation interference during policy updates by shuffling and batch sampling past samples. These samples serve as training inputs to optimize the Q network by minimizing the TD error, thereby guiding subsequent policy iteration and equilibrium optimization.
[0174] In the network training and parameter update steps, the following process is performed to achieve the minimax optimal strategy learning of the communication decision-maker and the security decision-maker in the core network under the continuous strategy space: (1) Batch experience sampling: Randomly sample interaction samples of batch size B from the experience replay pools of the communication decision-maker and the security decision-maker. Each sample is in the form of ; (2) Target Q value calculation: In each round of training, the target Q network of the communication party and the security party is used to calculate the maximum expected return in the next state; the expression of the expected return of the second communication decision-making party is ; The expected return of the second security decision-maker is expressed as ; (3) Loss function construction: In order to achieve the approximation of the target action value function by the Q network, the loss function of the second communication decision-maker as shown in the following expression (43) and the loss function of the second security decision-maker as shown in the following expression (44) are constructed: ;(43) ;(44) The function updates the current network parameters by gradient descent through backpropagation, realizing dynamic optimization of the strategy in the state space, and then continuously approaching the optimal game strategy equilibrium solution; (4) Network parameter optimization: Gradient update is performed on the main Q network of the second communication decision-maker and the second security decision-maker according to the communication party update rule and the security party update rule respectively; (5) Target network synchronization: To alleviate the problem of training instability, a target network soft update mechanism is introduced, which is based on and Update the model parameters; when When , the target network updates slowly, which helps stabilize the training; when , the target network quickly tracks the current strategy and improves the response speed. In addition, the target network synchronization operation can be performed after each round of training, or set to a fixed interval step periodic update. Through the soft update method, the target network gradually approaches the current strategy value function, which helps to strengthen the stable convergence of the learning process and the continuous improvement of the strategy optimality.
[0175] In the termination and policy convergence determination steps, a dual convergence determination mechanism based on both loss function changes and policy distribution fluctuations is designed to ensure that the game training process between the communicating and secure parties in the continuous policy space automatically terminates when a stable policy is reached. Specifically, the training termination criteria include the following two aspects: first, loss function convergence detection. This involves monitoring the Q-network loss functions of the communicating and secure parties during each training iteration through a sliding window. If the loss decreases below a preset threshold over K consecutive rounds, the Q-network is considered to have converged and no further training is required. Second, policy output stability detection is performed. This involves monitoring the changes in the action selection distributions (policy probability vectors) of the communicating and secure parties between consecutive iterations. If the KL divergence or Euclidean distance falls below a threshold, the policy output is considered stable and training can be terminated. Ultimately, when both convergence conditions are met, the system outputs the optimized communication and security policies as the game equilibrium solutions in the continuous policy space, which are used for the subsequent joint deployment and dynamic scheduling of secure communication policies.
[0176] Optionally, after obtaining the second optimized communication security policy based on the optimized communication policy and the optimized security policy, the method further includes: Convert the second optimized communication security strategy into system deployment parameters, including control instructions such as bandwidth allocation, power scheduling, encryption strength, and access control level of the core network, to achieve the optimal balance between communication performance and network security goals; Based on the second optimized communication security strategy, a game equilibrium value for characterizing the competitive joint utility of the internal competition game of the core network is calculated and output to evaluate the deployment effect of the secure communication strategy.
[0177] The policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention solves the problem of solving the minimax strategy optimization of the zero-sum game structure in the continuous strategy space based on the deep Q network, thereby optimizing the communication security strategy in the continuous variable space within the core network, realizing the intelligent coordination of dynamic strategy adjustment and game equilibrium deployment, and effectively improving the robustness and real-time performance of the communication security mechanism within the core network.
[0178] Based on the above embodiment, as an optional embodiment, the iterative execution of steps S20 to S40 to obtain the global optimal communication security policy of the wireless communication system includes: Iteratively executing steps S20 to S40 until the optimized communication delay division ratio converges, or until the first optimized communication security strategy and the second optimized communication security strategy are stable, thereby obtaining the first optimal communication security strategy and the second optimal communication security strategy; The global optimal communication security policy is determined based on the first optimal communication security policy and the second optimal communication security policy.
[0179] Specifically, in order to obtain the global optimal communication security strategy of the wireless communication system, steps S20 to S40 are iteratively executed, and the global iterative optimization of "external competition game - access network internal cooperation game / core network internal competition game - external competition game" is continuously performed until the optimized communication delay division ratio converges, or until the first optimized communication security strategy and the second optimized communication security strategy are stable, that is, the first optimal communication security strategy of the access network and the second optimal communication security strategy of the core network are obtained. By using the first optimal communication security strategy and the second optimal communication security strategy, the global optimal communication security strategy of the wireless communication system can be obtained.
[0180] The strategy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention can achieve dynamic game equilibrium between the access network and the core network by performing global iterative optimization of "external competition game - access network internal cooperation game / core network internal competition game - external competition game", and improve the overall synergistic optimal effect between the communication efficiency and security assurance capability of the wireless communication system.
[0181] Based on the above embodiment, as an optional embodiment, the security strength of the global optimal communication security policy is greater than the minimum security strength of the mobile communication service; the minimum security strength is determined based on the following method; Determine the current wireless channel environment threat index based on the abnormal behavior index of the current access requesting terminal and the frequency of unsuccessful matching between the access point and the terminal security mechanism per unit time; determining a risk index for the mobile communication service based on the sensitivity of the mobile communication service to security risks and the threat index of the current wireless channel environment; The minimum security strength is determined based on the risk index of the mobile communication service, the service fixed security strength level, and the maximum risk level that can be coped with by the service fixed security strength level.
[0182] Specifically, from the perspective of terminal behavior monitoring, the abnormal behavior index of the terminal currently requesting access is analyzed based on an intelligent algorithm in advance. Behavioral Abnormality Index The higher the value, the more challenges it faces in accessing the network, that is, configuring a higher security strength. From the perspective of security function log analysis, the frequency of unmatched access points and terminal security mechanisms failing to match successfully within a unit time is calculated. To analyze the number of abnormal interruptions, the more abnormal terminals (eavesdroppers / interceptors) in the wireless communication system, the worse the current wireless channel environment is, and a higher security strength needs to be configured to cope with environmental risks.
[0183] Based on the abnormal behavior index of the current request access terminal The frequency of unmatched access points and terminal security mechanisms failing to match successfully per unit time , you can determine the current wireless channel environment threat index .
[0184] Furthermore, according to the current wireless channel environment threat index The risk index of mobile communication services can be obtained by combining the sensitivity of mobile communication services to security risks. The sensitivity of mobile communication services to security risks is a service sensitivity coefficient preset by the system or obtained through data-driven learning.
[0185] The minimum security strength is then determined using the risk index of the mobile communication service, the fixed security strength level of the service, and the maximum risk level that the fixed security strength level of the service can cope with.
[0186] It is understandable that the minimum security strength, as a unified security constraint premise for system strategy optimization, is effective by default in various external games, internal games and deep reinforcement learning modeling.
[0187] In one embodiment, the expression for determining the current wireless channel environment threat index based on the abnormal behavior index of the terminal currently requesting access and the frequency of unsuccessful matching between the access point and the terminal security mechanism per unit time is as follows: ;(45) in, is the current wireless channel environment threat index; is the behavioral abnormality index; The frequency is not matched.
[0188] In one embodiment, based on the current wireless channel environment threat index and the sensitivity of the mobile communication service to security risks, an expression for determining the risk index of the mobile communication service is as follows: ;(46) in, is the risk index of the i-th mobile communication service; is the sensitivity of the i-th mobile communication service to security risks; is the current wireless channel environment threat index.
[0189] In one embodiment, based on the risk index of the mobile communication service, the fixed security strength level of the service, and the maximum risk level that the fixed security strength level of the service can cope with, an expression for determining the minimum security strength is as follows: ; (47) in, is the minimum security strength of the i-th mobile communication service; is the risk index of the i-th mobile communication service; A fixed security strength level for the service of the i-th mobile communication service; The maximum risk level that can be handled by the fixed security strength level of the i-th mobile communication service.
[0190] The present invention provides a policy optimization method for a high-security wireless communication system based on hierarchical game. By analyzing the abnormal behavior indications and mismatch frequencies from the two perspectives of terminal behavior monitoring and security function log analysis, the time-varying wireless channel environment threat index is determined. In combination with the service characteristics and the initial security indicators of the service, the minimum security level required for each service is dynamically adjusted, thereby ensuring the security and feasibility of the communication security policy optimization of the wireless communication system.
[0191] Based on the above embodiment, as an optional embodiment, the following is further included: S60: Determine a target security function module to be called according to a global optimal communication security policy; call the target security function module from an extensible security function resource pool to implement secure communication of the wireless communication system.
[0192] Specifically, the scalable security function resource pool is designed based on the implementation of 5G network function virtualization, which virtualizes each security function into an independent module. After completing the global game optimization of the wireless communication system to obtain the global optimal communication security strategy, several target security function modules to be called are determined according to the global optimal communication security strategy, and then the target security function modules are called from the scalable security function resource pool to realize secure communication of the wireless communication system.
[0193] Among them, the security function modules in the scalable security function resource pool include but are not limited to wireless data stream encryption module, integrity protection module, device trusted authentication module and privacy protection module.
[0194] The policy optimization method for a high-security wireless communication system based on hierarchical game provided by the present invention builds an extensible security function resource pool based on the concept of 5G network function virtualization (NFV), modularizes security functions such as trusted authentication, data encryption, and privacy protection into virtual resources, and provides plug-and-play flexibility, allowing users to dynamically adjust or add security modules according to actual needs, thereby enhancing the system's adaptability to dynamic changes in the network environment, ensuring continuous, flexible, and efficient security protection, and realizing dynamic security protection of an extensible security function resource pool.
[0195] Figure 2 Schematic diagram of the structure of the strategy optimization device of the high-security wireless communication system based on hierarchical game provided by the present invention. Figure 2 As shown, the policy optimization device of the high-security wireless communication system based on hierarchical game includes a threat situation intelligent perception module, a security level dynamic control module, an extensible security function resource pool and an adaptive security function orchestration module.
[0196] The threat situation intelligent perception module is used to determine the current wireless channel threat index based on the abnormal behavior index of the terminal currently requesting access and the frequency of unsuccessful matching between the access point and the terminal security mechanism per unit time. Through comprehensive analysis of terminal behavior monitoring and security function logs, the threat index in the wireless channel environment is evaluated in real time, aiming to provide effective decision-making reference for the dynamic security level control module. The dynamic security level control module is configured to determine the risk index of a mobile communication service based on the service's sensitivity to security risks and the current wireless channel threat index. This module dynamically controls the minimum security level required for each service by combining service characteristics, initial security indicators, and the time-varying wireless channel threat index.
[0197] The scalable security function resource pool, based on the concept of 5G network function virtualization, virtualizes various security functions into independent modules. The system can configure or add security functions to each terminal as needed, enabling plug-and-play security service deployment. This simplifies the expansion and orchestration of security functions and provides flexible and efficiently scalable security resources for high-security mobile communication systems. Through this virtualized architecture, the security function resource pool can be flexibly configured based on network environments and business needs, meeting the personalized security needs of different users and application scenarios.
[0198] The adaptive security function orchestration module is configured to execute steps S10 to S60. By introducing a hierarchical game mechanism, including external game (between the access network and the core network) and internal game (between communication and security decision-makers within the access network and the core network), it aims to achieve overall optimization and balance between system security and communication efficiency.
[0199] Figure 3 This is a schematic diagram of the hierarchical game mechanism architecture provided by the present invention, such as Figure 3 As shown, the external game constructs a competitive game model between the access network and the core network. By dynamically allocating the maximum tolerable end-to-end delay and determining the delay budgets for each network, the global optimization of communication performance and security protection is achieved. In the internal game, a cooperative game is adopted between the communication decision-maker and the security decision-maker within the access network. The communication decision-maker primarily optimizes wireless access performance, including optimal access point selection, optimal transmit power setting, and optimal bandwidth allocation. The security decision-maker focuses on implementing wireless security policies and providing security protection through the wireless channel, including wireless data flow encryption, data integrity protection, device authentication, and privacy protection. Furthermore, the two parties' objectives are in a competitive relationship: the communication decision-maker seeks to minimize communication delay, while the security decision-maker's security policies typically increase communication delay. Through a cooperative game mechanism, both parties dynamically maximize wireless security protection capabilities while ensuring that communication delay meets minimum QoS requirements, achieving an optimal balance between communication performance and security benefits. A competitive game is adopted between the communication decision-makers and security decision-makers within the core network. The communication decision-makers focus on optimizing link transmission delay to reduce the communication overhead caused by security measures; the security decision-makers focus on data security protection in wired networks, mainly considering the transmission issues of valid data and encrypted data (such as data encryption header length, key strength, etc.). The two parties compete for the limited total core network delay and ultimately strike a balance between minimizing link transmission delay and maximizing data security encryption strength.
[0200] Overall, the present invention proposes a hierarchical game mechanism, constructing a competitive game model between the access network and the core network, and further establishing a cooperative game model between the communication decision-makers and the security decision-makers within the access network, and a competitive game model between the communication decision-makers and the security decision-makers within the core network, thereby dynamically optimizing the allocation of communication resources and the combination of security policies, ensuring that the system can adapt to the real-time changing network environment and security threats. At the same time, the present invention integrates mathematical optimization methods with heuristic optimization algorithms based on deep reinforcement learning, effectively realizing intelligent dynamic scheduling of resources and adaptive optimization of security policies, enabling wireless communication systems to perceive network security situations in real time, actively respond and dynamically adjust defense strategies; by adopting Nash equilibrium analysis in the game process, based on multi-party interactive optimization, determining the security level of each service, and rationally allocating communication resources and security algorithm strength, ensuring that high-risk services can have priority access to higher-intensity security protection and more abundant communication resources. Ultimately, under the dynamic protection of security levels, the system achieves the optimal coordination of communication efficiency and protection strength, providing stronger flexibility and reliability support for the operation of wireless systems in high-security scenarios, so that the security protection capabilities and communication stability of high-security wireless communication systems in sensitive fields such as military, finance, and medical care will be significantly improved, which will help promote the development of high-security wireless communication technology towards intelligence and autonomy.
[0201] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A strategy optimization method for a high-security wireless communication system based on hierarchical game, characterized in that: include: S10. Determine a current communication delay division ratio of the wireless communication system, a first current communication security policy of the access network, and a second current communication security policy of the core network; S20. With the goal of optimizing the overall communication security benefits of the access network and the core network, iteratively optimize the current communication delay division ratio to obtain an optimized communication delay division ratio; S30. Optimize the first current communication security policy and the second current communication security policy under the optimized communication delay division ratio to obtain a first optimized communication security policy for the access network and a second optimized communication security policy for the core network; S40, determining a new current communication delay division ratio, a new first current communication security policy, and a new second current communication security policy based on the optimized communication delay division ratio, the first optimized communication security policy, and the second optimized communication security policy; S50. Iteratively execute steps S20 to S40 to obtain a global optimal communication security policy for the wireless communication system; the global optimal communication security policy includes a first optimal communication security policy for the access network and a second optimal communication security policy for the core network.
2. The strategy optimization method for a high-security wireless communication system based on hierarchical game according to claim 1, characterized in that: Step S20 includes: Under the first KKT condition, with the goal of optimizing the overall communication security benefit of the access network and the core network, iteratively optimize the current communication delay partition ratio based on a gradient descent method and a first Lagrangian objective function to obtain the optimized communication delay partition ratio; The first Lagrangian objective function includes an external joint utility function and an upper limit constraint on communication delay allocation; The external joint utility function includes an access network utility function and a core network utility function; the access network utility function is determined based on a communication policy benefit and a communication policy overhead; the core network utility function is determined based on a security policy benefit and a security policy overhead; the communication policy benefit and the communication policy overhead are determined based on a first current communication security policy and a current communication delay division ratio; the security policy benefit and the security policy overhead are determined based on a second current communication security policy and a current communication delay division ratio; The communication delay allocation upper limit constraint is determined based on the current total budget of end-to-end delay, the adjustable upper bound of the total budget of end-to-end communication delay, and the resource sensitivity of the adjustable upper bound of the total budget of end-to-end communication delay; the Lagrangian multiplier of the first Lagrangian objective function is the resource sensitivity.
3. The strategy optimization method for a high-security wireless communication system based on hierarchical game according to claim 1, characterized in that: Step S30 includes: S31, with the goal of optimizing the cooperative joint utility between the first communication decision-maker and the first security decision-maker of the access network under the optimized communication delay division ratio, iteratively optimize the first current communication security policy to obtain a first optimized communication security policy; S32. With the goal of optimizing the competitive joint utility of the second communication decision-maker and the second security decision-maker of the core network under the optimized communication delay division ratio, the second current communication security policy is iteratively solved to obtain a second optimized communication security policy.
4. The strategy optimization method for a high-security wireless communication system based on hierarchical game according to claim 3, characterized in that: The first current communication security policy includes the communication policies of the first communication decision-maker and the security policies of the first security decision-maker; the cooperative joint utility includes the communication decision-maker utility and the security decision-maker utility; step S31 includes: Under a second KKT condition, with the goal of optimizing the cooperative joint utility of the access network, the first current communication security policy is iteratively optimized based on a policy update rule and a second Lagrangian objective function to obtain the first optimized communication security policy; the policy update rule includes sub-update rules for each communication policy and each security policy in the first current communication security policy; The second Lagrangian objective function includes a cooperative joint utility function, an access network communication delay budget constraint, a bandwidth resource constraint, and a security constraint; The cooperative joint utility function is determined based on the communication decision-maker utility function and the security decision-maker utility function; the communication decision-maker utility function is determined based on the policy benefits of each communication policy of the first communication decision-maker and the delay overhead of each security policy of the first security decision-maker; the security decision-maker utility function is determined based on the anti-attack capability of each security policy of the first security decision-maker and the security synergy benefit of each communication policy of the first communication decision-maker; The access network communication delay budget constraint is determined based on the optimized communication delay division ratio and the access network delay budget sensitivity; the bandwidth resource constraint is determined based on the access network allocated bandwidth, the total system available bandwidth and the bandwidth resource sensitivity in the first current communication strategy; the security constraint is determined based on the minimum access network security utility, the security decision-maker utility and the security sensitivity; the Lagrangian multiplier of the second Lagrangian objective function includes the access network delay budget sensitivity, the bandwidth resource sensitivity and the security sensitivity.
5. The strategy optimization method for a high-security wireless communication system based on hierarchical game according to claim 4, characterized in that: The step of iteratively optimizing the first current communication security policy based on a policy update rule and a second Lagrangian objective function with the goal of optimizing the cooperative joint utility of the access network to obtain the first optimized communication security policy includes: constructing a first input vector based on the first current communication security policy; Inputting the first input vector into a first policy update model to obtain a first optimized communication security policy output by the first policy update model; the first policy update model is implemented based on a deep Q network; The state space of the first policy update model is determined based on the communication strategies and security strategies in the first current communication security strategy; the action space of the first policy update model is determined based on the policy update rule; and the reward function of the first policy update model is determined based on the cooperative joint utility function in the second Lagrangian objective function.
6. The strategy optimization method for a high-security wireless communication system based on hierarchical game according to claim 3, characterized in that: In the case where the second current communication security policy is a discrete policy, step S32 includes: Based on the second current communication security policy, defining a communication hybrid policy set of a second communication decision-maker of the core network and a security hybrid policy set of a second security decision-maker of the core network; defining a first hybrid strategy probability distribution variable corresponding to the communication hybrid strategy set and a second hybrid strategy probability distribution variable corresponding to the security hybrid strategy set; Constructing a game matrix; each element in the game matrix is a loss function value obtained by substituting a communication hybrid strategy from the communication hybrid strategy set and a security hybrid strategy from the security hybrid strategy set into a pre-constructed communication party loss function; the communication party loss function is determined based on the core network end-to-end delay and the security policy communication overhead; Based on the optimized communication delay partition ratio, the game matrix, and the first mixed strategy probability distribution variable, with the goal of minimizing the upper limit of communication expected loss, constructing the original linear programming model of the second communication decision-maker; Based on the optimized communication delay partition ratio, the game matrix, and the second mixed strategy probability distribution variable, a dual linear programming model of the second security decision-maker is constructed with the goal of maximizing the lower bound of the expected security benefit; Calling a standard linear programming solver to jointly solve the original linear programming model and the dual linear programming model to obtain an optimal solution pair of the hybrid strategy probability distribution with the optimal competitive joint utility; The second optimized communication security strategy is determined based on the optimal solution pair of the hybrid strategy probability distribution.
7. The strategy optimization method for a high-security wireless communication system based on hierarchical game according to claim 3, characterized in that: In the case where the second current communication security policy is a continuous policy, step S32 includes: constructing a second input vector based on the communication policy in the second current communication security policy; Inputting the second input vector into a second policy update model to obtain an optimized communication policy output by the second policy update model; the second policy update model is implemented based on a deep Q network; constructing a third input vector based on the security policy in the second current communication security policy; Inputting the third input vector into a third policy update model to obtain an optimized security policy output by the third policy update model; the third policy update model is implemented based on a deep Q network; Based on the optimized communication strategy and the optimized security strategy, obtaining the second optimized communication security strategy; The state spaces of the second policy update model and the third policy update model are both determined based on each communication policy and each security policy in the second current communication security policy; The action space of the second policy update model is determined based on the communication policy in the second current communication security policy; the reward function of the second policy update model is determined based on the core network end-to-end delay and the security policy communication overhead; The action space of the third policy update model is determined based on the security policy in the second current communication security policy; the reward function of the third policy update model is determined by the security policy benefit and the security policy encryption processing delay.
8. The strategy optimization method for a high-security wireless communication system based on hierarchical game according to claim 1, characterized in that: The iterative execution of steps S20 to S40 to obtain a global optimal communication security policy for the wireless communication system includes: Iteratively executing steps S20 to S40 until the optimized communication delay division ratio converges, or until the first optimized communication security strategy and the second optimized communication security strategy are stable, thereby obtaining the first optimal communication security strategy and the second optimal communication security strategy; The global optimal communication security policy is determined based on the first optimal communication security policy and the second optimal communication security policy.
9. The strategy optimization method for a high-security wireless communication system based on hierarchical game according to claim 8, characterized in that: The security strength of the global optimal communication security policy is greater than the minimum security strength of the mobile communication service; the minimum security strength is determined based on the following method; Determine the current wireless channel environment threat index based on the abnormal behavior index of the current access requesting terminal and the frequency of unsuccessful matching between the access point and the terminal security mechanism per unit time; determining a risk index for the mobile communication service based on the sensitivity of the mobile communication service to security risks and the threat index of the current wireless channel environment; The minimum security strength is determined based on the risk index of the mobile communication service, the service fixed security strength level, and the maximum risk level that can be coped with by the service fixed security strength level.
10. The strategy optimization method for a high-security wireless communication system based on hierarchical game according to claim 1, characterized in that: Also includes: S60, determining a target security function module to be called according to a global optimal communication security strategy; The target security function module is called from the scalable security function resource pool to implement secure communication of the wireless communication system.