System security management method
By combining system security management methods with reinforcement learning and optimal control theory, user permissions are dynamically adjusted, and the game theory model is used to optimize the interaction between the system and users, the problems of lagging permission adjustment and imbalance in the existing technology are solved, and efficient security management and resource utilization are achieved.
Patent Information
- Application Number
- CN202510275523.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The prior art cannot adjust permissions in time when facing dynamic and complex user behavior, resulting in increased security risks and traditional security protection strategies are difficult to balance security and resource utilization efficiency.
A system security management method combining reinforcement learning and optimal control theory is adopted to monitor user behavior and risk assessment in real time, dynamically adjust user permissions, and use game theory models to optimize the interaction between the system and the user to achieve a balance between security and resource utilization.
It realizes rapid response to abnormal behaviors and dynamic adjustment of permissions, improves the flexibility and response speed of the system, and ensures the best balance between security and resource utilization.
Smart Images

Figure CN120105460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security, and in particular to a system security management method. Background Art
[0002] Existing permission management methods usually rely on static rules and role assignments, and the system sets each user's access rights based on fixed rules. This method can play a basic role in simple scenarios, but it has obvious shortcomings when faced with dynamic and complex user behaviors. Since it is impossible to make corresponding permission adjustments based on changes in user behavior in real time, traditional methods often result in the system being unable to take timely action when abnormal behavior occurs, increasing potential security risks.
[0003] The risk assessment mechanism in the prior art usually relies on historical data and manually set thresholds, and lacks real-time monitoring and adaptive capabilities for user behavior. Since this method relies too much on preset rules and static data, the system is often unable to respond quickly to emerging security threats. For example, when a user's behavior changes unexpectedly, such as accessing highly sensitive data or logging in from an uncommon location, the system may not be able to adjust permissions in time, resulting in security vulnerabilities. The present invention combines real-time behavior monitoring with dynamic risk assessment to achieve rapid response to abnormal behavior and dynamic adjustment of permissions, solving the problem of delayed response of traditional methods in the face of emerging threats.
[0004] Traditional security protection strategies focus too much on system security and ignore the balance with resource utilization. Security is usually guaranteed by strict permission restrictions. However, this strategy may lead to waste of resources and reduced user experience, especially in a multi-user environment. It is difficult for existing systems to balance security and resource utilization efficiency, especially when it is necessary to ensure resource access for a large number of users. The present invention can dynamically optimize resource allocation while ensuring system security through a game theory model and an adaptive feedback mechanism, thereby achieving a balance between security and resource utilization. Therefore, the present invention proposes a system security management method to solve the deficiencies of the prior art. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides a system security management method, which solves the problems of lack of dynamic adjustment, lagging risk assessment, and imbalance between security and resource utilization in traditional authority management; by introducing reinforcement learning and optimal control theory, the present invention can dynamically adjust user permissions according to real-time risk assessment results, respond to abnormal behaviors in a timely manner, and improve system flexibility and response speed; in addition, a game theory model is used to optimize the interaction between the system and the user, balance security and resource utilization, and ensure that the system does not excessively restrict user resource access while providing high security, thereby achieving the best balance between system security and resource utilization.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a system security management method, comprising the following steps: Monitor user behavior data in real time, assess its security risks, and provide assessment results; Based on the evaluation results, the user's access rights are dynamically adjusted through the reinforcement learning model; Based on optimal control theory, optimize the user's access rights adjustment strategy to maximize the system's security and resource utilization; use game theory models to optimize the behavioral interaction between the system and users to solve the optimal security protection strategy; Through the adaptive feedback mechanism, key parameters in the reinforcement learning model, optimal control model and game theory model are adjusted according to the real-time feedback data of the system operation to improve the responsiveness and flexibility of the system.
[0007] Preferably, the step of real-time monitoring of user behavior data includes: Obtain the user's login information, operation frequency, and access resource type and behavior characteristics; Combine users’ historical behavior data with real-time behavior data to assess potential security risks; The user behavior data includes but is not limited to: User's login information; Operating frequency; The type of resource being accessed; User's historical behavior data; Other relevant behavioral characteristics; The user behavior data is protected through encryption, desensitization and other technologies to ensure the privacy and compliance of the data.
[0008] Preferably, the reinforcement learning model adjusts access rights based on the user's behavioral characteristics and real-time access requests by modeling the state space and action space of the user's behavior.
[0009] Preferably, the reinforcement learning model uses a Q-learning algorithm to calculate the expected reward of each user behavior and adjust the user access rights in real time based on the reward.
[0010] Preferably, the optimal control theory optimizes the access permission adjustment strategy by minimizing the cost function to balance system security and resource utilization, and the cost function includes: Where: J(t) is the total cost function; S(t) is the security score of the system at time t; P(t) is the degree of adjustment of the system's permissions to the user; w 1 and w 2 is the weight coefficient.
[0011] Preferably, the step of optimizing the user authority adjustment strategy by the optimal control theory includes: Set the system state equation to represent the dynamic changes of user behavior and access requests; Real-time adjustment of user access rights is achieved by calculating the optimal control input.
[0012] Preferably, the game theory model optimizes the protection and resource access game between the system and the user by solving the Nash equilibrium, so that the behavior between the system and the user is balanced.
[0013] Preferably, the security protection strategy optimized by the game theory model includes: The system maximizes security and users maximize access to resources; The strategies obtained through game analysis ensure security while not placing excessive restrictions on the user's operating experience.
[0014] Preferably, the adaptive feedback mechanism includes: Adjust parameters in reinforcement learning, optimal control, and game theory models based on real-time risk assessment and behavioral data; Combine immediate feedback with historical behavior analysis to optimize user permission adjustment strategies.
[0015] The present invention also provides a system security management system, comprising: Behavior monitoring module, used to monitor user behavior characteristic data in real time and conduct security risk assessment; The permission adjustment module is used to dynamically adjust user access rights based on reinforcement learning and optimal control models; Game theory model module, used to optimize the interaction between the system and users and solve the optimal protection strategy; Adaptive feedback module, which is used to adjust model parameters according to real-time feedback to improve the responsiveness and flexibility of the system.
[0016] The present invention provides a system security management method, which has the following beneficial effects: 1. The present invention adopts a technical solution based on the combination of reinforcement learning and optimal control theory, which achieves the effect of dynamically adjusting user permissions in an intelligent way, maximizing system security and optimizing resource utilization; compared with the static permission management solution in the prior art, the present invention can effectively respond to ever-changing security threats through real-time feedback and learning, and solves the problem that traditional methods cannot flexibly adjust permissions.
[0017] 2. The present invention adopts a game theory model to optimize the behavioral interaction between the system and the user, and solves the optimal security protection strategy, thereby achieving the effect of reasonably meeting user needs while ensuring system security; compared with the one-size-fits-all permission control method in the existing solution, the present invention uses the game theory game strategy to enable the system and users to achieve the best balance between security and resource access, solving the problem of excessive restrictions or resource waste.
[0018] 3. The present invention uses an adaptive feedback mechanism to adjust model parameters in real time according to system operation feedback, thereby achieving the effect of improving system flexibility and response speed; different from the fixed rule model in traditional technology, the present invention can quickly adapt to changes in the external environment and user behavior, reducing the problem of delayed response in traditional solutions, and ensuring that the system can be quickly adjusted in the face of changing needs.
[0019] 4. The present invention achieves the effect of accurate assessment and real-time adjustment of user behavior through a behavior monitoring module and a real-time risk assessment mechanism, combined with dynamic adjustment and reinforcement learning models; compared with the technical solutions in the prior art that only rely on static rules for security assessment, the present invention solves the problem of untimely detection of potential risks through real-time collection and analysis of behavioral data, and effectively improves the security protection capabilities of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION
[0021] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] See also Figure 1 , an embodiment of the present invention provides a system security management method, comprising the following steps: S1. Monitor user behavior data in real time, assess its security risks, and give assessment results; In this embodiment, step S1 involves real-time monitoring of user behavior data and performing security risk assessment on it; this step occupies a fundamental position in the entire system security management method and provides a decision-making basis for subsequent permission adjustments; real-time monitoring of user behavior and assessment of potential security risks enable the system to respond to abnormal behavior in a timely manner and ensure the dynamic security of the system.
[0023] In this embodiment, the core function of the behavior monitoring module is to collect and analyze user behavior data, including user login information, operation frequency, resource access type, etc.; the system forms a user behavior feature model by analyzing these data; based on this model, the system can determine whether the user behavior is in line with expectations, and score each behavior according to preset security risk assessment rules.
[0024] During the implementation process, the system first collects user behavior data from various data sources; these data include but are not limited to: User login information: login time, device information, geographic location, etc.; User operation frequency: access frequency, access duration, type of interactive behavior, etc. Resource access type: the type of resource accessed (such as sensitive data, public data, etc.) and the access method (such as read, modify, etc.); Access behavior characteristics: such as whether the user logs in through a common device, whether the user frequently switches geographical locations, etc. This information is collected in real time and input into the system's behavior analysis module as feature vectors.
[0025] The system evaluates user behavior based on historical behavior data and real-time behavior data. To achieve this goal, the system first maps the behavioral characteristics of each user into the state space. Then, it determines whether the user behavior is normal by calculating the degree of behavioral abnormality.
[0026] The system's behavioral risk assessment formula is as follows: R(t)=α·S(t)+β·F(t); Among them, R(t) represents the risk score at time t, reflecting the degree of abnormality of user behavior; S(t) represents the security score of the system, evaluating the overall security status of the system; F(t) represents the abnormality of user behavior, quantifying whether the user behavior deviates from the normal mode; α and β are weight coefficients, which respectively adjust the impact of security score and abnormality on the final risk score.
[0027] Specifically, S(t) and F(t) are calculated as follows: Safety score S(t): S(t) = f(operation frequency, frequency of access to sensitive resources, consistency of behavior and history); Among them, S(t) reflects the normality of user behavior. Generally, users with low operation frequency, low access frequency to sensitive resources and consistent behavior with historical behavior are considered to have safer behavior. f(·) in this formula represents a function trained by historical data and pattern recognition algorithm, which is used to quantify the difference between normal behavior and abnormal behavior. Abnormality score F(t): F(t) = g(abnormal login IP, behavioral pattern deviation, abnormal time period); Among them, F(t) represents the degree of abnormality of user behavior. For example, if the user's login IP changes suddenly, or the user's access period no longer conforms to his historical behavior pattern, it may lead to a large F(t) value. g(·) is a function based on behavior pattern recognition, which can calculate the degree of abnormality of user behavior.
[0028] In this embodiment, the system sets a risk assessment threshold R th ; If a user's risk score R(t) exceeds the threshold, the user's behavior will be marked as abnormal and trigger the subsequent emergency response mechanism; the threshold R th It can be adjusted according to the actual application scenario and set according to the sensitivity of user access and the degree of abnormal behavior.
[0029] For example, when R(t)>R th , the system can respond in the following ways: Triggering permission restrictions: If abnormal user behavior is detected, the system may restrict their access rights; Triggering secondary authentication: If a user accesses sensitive resources and behaves abnormally, the system can require the user to perform additional identity authentication (such as secondary authentication, SMS verification, etc.); Alert mechanism: The system will immediately generate an alert and notify the administrator for review.
[0030] For example, suppose a user usually logs in through the company's internal network and common devices, but one day the user logs in from a different location and attempts to access highly sensitive data; the system will evaluate it through the following steps: Monitor user behavior: collect user login information, access resources and other data.
[0031] Calculate the security score S(t): Assuming that the user's operation frequency is normal and consistent with historical behavior, S(t) may be a higher value; Calculate the abnormality score F(t): Since the user logged in from a different location and accessed sensitive data, the system evaluates the abnormality of this behavior, and F(t) is high; Calculate the risk score R(t): Combine S(t) and F(t) to calculate the total risk score. If R(t) exceeds the set threshold R th , then the behavior is considered to be risky.
[0032] If the risk score is higher than the set threshold, the system may restrict the user's access and request authentication; the administrator will also receive an alert and review the behavior.
[0033] S2. Dynamically adjust the user's access rights through the reinforcement learning model based on the evaluation results; In step S2, the system dynamically adjusts the user's access rights through a reinforcement learning model based on the evaluation results obtained in step S1. This process is completed on the basis of real-time monitoring of user behavior and assessment of security risks. It aims to optimize access control policies through adaptive learning to ensure that the system can respond promptly to changing security threats.
[0034] In this embodiment, the core of step S2 is the reinforcement learning model, which dynamically adjusts the user's access rights based on the user's behavior data and the system's security assessment results; in the aforementioned step S1, the system has calculated the risk score R(t) based on the user's operating behavior and behavior characteristics, and the score reflects the degree of abnormality and potential risks of the user's behavior; in this step, the reinforcement learning model will automatically adjust the user's permissions based on this score to ensure the security of the system.
[0035] Reinforcement Learning (RL) is an adaptive learning algorithm that optimizes strategies through the interaction between the environment and actions. In this embodiment, the reinforcement learning model adjusts user access rights and optimizes the security of the system by continuously observing user behavior and calculating the risk score R(t).
[0036] The reinforcement learning model in the system can be divided into three main parts: State space S(t): represents the complete state of the system at a certain moment; Action space A(t): represents the access rights adjustment strategies that the system can adopt; Reward function R(t): used to evaluate the performance of the system after adopting a certain strategy.
[0037] The state space S(t) represents the complete state of the system at time t, which is usually composed of multiple behavioral features that reflect the user's behavior and the deviation of the behavior from historical data. The state space usually includes the following: User behavior characteristics: such as operation frequency, login device, accessed resource type, operation duration, etc. Safety score S(t): evaluates the normality of the current user behavior. The closer to normal, the higher the score; Abnormality score F(t): reflects the abnormality of user behavior. The higher the score, the more abnormal the behavior. Therefore, the state space S(t) can be expressed as: S(t)=(B 1 (t),B 2 (t),…,B N (t), S(t), F(t)); Among them, B i (t) represents the behavior characteristics of the i-th user at time t, S(t) is the security score, F(t) is the anomaly score, and N is the number of users in the system.
[0038] The action space A(t) represents all the permission adjustment operations that the system can select; each action corresponds to the adjustment of user permissions; in this embodiment, the action space of the system includes: Permission reduction: When a user's risk score is high, the system will tighten their access rights and restrict their access to certain resources; Permission expansion: When a user's behavior is normal and there is no abnormal risk, the system will expand their permissions and allow them to access more resources; Maintain existing permissions: For users with no abnormal behavior and low risk scores, the system will maintain their existing permissions; Therefore, the action space A(t) can be defined as: A(t)=(P 1 (t),P 2 (t),…,P N (t)); Among them, P i (t) represents the access rights of the ith user at time t, P i (t)∈[0,1], that is, the larger the permission value, the higher the degree of access allowed. N is the number of these elements. There are N such elements or variables in the system.
[0039] The reward function is the core of the reinforcement learning model, which is used to evaluate the feedback of the system after taking a certain action. The reward function will combine the security of the system and the effect of user permission adjustment to give a corresponding reward value. The goal of the system is to maximize the security score while trying to avoid excessive restrictions on permissions and reduce resource waste. The reward function in this embodiment is as follows: R(t) = α·S(t) - β·P(t); Among them: S(t) is the security score of the system, which reflects the security of user behavior in the current state; P(t) is the degree of user authority adjustment, which indicates the size of shrinking or expanding authority; α and β are weight coefficients, α controls the impact of the security score on the reward, and β controls the penalty effect of the authority adjustment degree; Reinforcement learning uses the Q-learning algorithm to update the Q value according to the reward function. The Q value represents the expected reward for taking a certain action in the current state. The formula for updating the Q value is: Where: Q(s,a) is the Q value of taking action a in state s; It represents the expected value of a random variable (or random process); R(t) is the immediate reward; γ is the discount factor, which represents the weight of future rewards, γ∈[0,1]; Indicates that in the next state s ′ Next, take the best action a ′ The maximum expected reward when By updating the Q value, the system can gradually optimize the permission adjustment strategy and select the action that is most beneficial to the system security.
[0040] The reinforcement learning model conducts online learning based on user behavior data and evaluation results, that is, it continuously updates the Q value based on the latest data obtained in real time. When the system detects that the user's behavior has changed, the Q value will be adjusted according to the new reward feedback, thereby optimizing the permission decision. For example, when the system finds that a user frequently accesses highly sensitive resources, his or her permissions will be reduced; and when the system evaluates that the user's behavior is normal, the permissions will be expanded.
[0041] Over time, the system gradually improves its understanding of user behavior through continuous learning and optimization, making permission adjustment strategies more precise.
[0042] For example, suppose a user usually works in the company's intranet and frequently accesses low-sensitivity resources; however, one day the user suddenly logs in from an external IP and attempts to access high-sensitivity resources; the system evaluates step S1 and calculates that the user's risk score R(t) is high and exceeds the set threshold; at this time, the system enters step S2 and dynamically adjusts the user's access rights based on the risk score and reward function.
[0043] Through the Q-learning algorithm, the system evaluates the user's behavior and decides to reduce his access rights, minimizing his access to highly sensitive resources; as the user's behavior changes further, the system will continue to monitor his risk score and continue to adjust his permissions based on the new score to ensure a balance between security and resource utilization.
[0044] S3. Based on optimal control theory, optimize the user's access rights adjustment strategy to maximize the system's security and resource utilization; In step S3, the system optimizes the user permission adjustment strategy through optimal control theory; the main purpose of this step is to implement an optimal access control strategy in the system to ensure a balance between security and resource utilization; by considering the user's risk score, access behavior, and system resource usage, the optimal control model helps the system dynamically adjust access rights to ensure the security and efficiency of the system.
[0045] In this embodiment, optimal control theory is used to optimize the access permission adjustment strategy. The specific implementation includes establishing a cost function, calculating the balance between security and resource consumption, and dynamically adjusting the access permissions of each user according to the optimization results. Through the optimal control algorithm, the system can respond to changing security threats and resource requirements in real time, so as to optimize the utilization of system resources while ensuring high security.
[0046] The core goal of the optimal control model is to minimize the cost function; the cost function is used to measure the performance of the system within a given time, and its design takes into account the trade-off between system security and resource utilization; specifically, the cost function needs to be comprehensively evaluated based on the current security status of the system and the user's access rights, so as to derive the optimal permission adjustment strategy.
[0047] The cost function J(t) is defined as: Where: J(t) is the total cost function, which represents the cumulative cost of the entire optimization process; S(t) is the security score of the system at time t, which represents the security status of the current system. The higher the score, the more secure the system; P(t) is the degree of adjustment of the system's permissions to users, which represents the control over resources. The higher the score, the greater the degree of permissions contraction; w 1 and w 2 is a weight coefficient used to balance the security score and resource consumption caused by permission adjustment; in practical applications, w 1 When w is larger, the system tends to improve security; 2 In larger cases, the system will try to avoid excessive restrictions on user permissions and optimize the use of resources.
[0048] In general, the goal of the cost function is to minimize resource consumption and over-contraction of permission adjustments while maximizing the security of the system.
[0049] In order to enable the optimal control algorithm to accurately perform permission optimization, the system defines a state equation that describes the relationship between user behavior and permission adjustment. The state equation establishes the dynamic relationship between user behavior, risk assessment and permission adjustment through a mathematical model. The state equation is as follows: x(t+1)=A·x(t)+B·u(t); Among them: x(t+1) is the state vector of the system at the next time t+1; x(t) is the system state vector, which contains the behavior characteristics, risk assessment results and current permission status of all users; A is the state transfer matrix, which represents the relationship between user behavior and system permission status; B is the control matrix, which represents the influence of control input on system status; u(t) is the control input, which represents the access rights adjusted by the system at time t.
[0050] The state equation helps the system calculate the optimal permission adjustment strategy based on the current state by describing the dynamic changes of the system.
[0051] To solve the optimal control problem, the system uses the linear quadratic (LQR) control method to obtain the optimal authority adjustment strategy by minimizing the cost function J(t); the LQR control algorithm calculates the optimal authority adjustment input u(t) at each time t by solving an optimal gain matrix K.
[0052] The optimal control input calculation formula is: u(t)=-Kx(t); Among them: u(t) is the optimal control input, which represents the optimal authority adjustment strategy under state x(t); K is the optimal gain matrix, which is calculated by the LQR control algorithm and is used to determine how to adjust the authority under different states; x(t) is the system state vector, which contains the behavior characteristics, risk scores and other data at the current time t.
[0053] Through the LQR control algorithm, the system can calculate the optimal permission adjustment strategy according to the current state at each moment, and adjust the user's access rights according to the optimal gain matrix K.
[0054] During the implementation process, the system adjusts permissions in real time based on user behavior data, risk assessment results and optimal control strategies; whenever the system receives new user behavior data, the optimal control algorithm will recalculate the permission adjustment strategy based on the current state equation and cost function to ensure the optimal balance between system security and resource utilization.
[0055] As an option, during the implementation process, the system can also introduce a feedback mechanism; that is, every time the system adjusts user permissions, it will recalculate the cost function and adjust the optimal control strategy based on the impact of the new permission adjustment on system security and resource utilization efficiency; through this real-time feedback, the system can continuously optimize the permission control strategy based on the ever-changing environment and behavioral data.
[0056] Suppose a user mainly accesses some low-sensitivity resources in daily work, but one day the user suddenly accesses a large number of high-sensitivity resources and logs in from different locations; the system evaluates through step S1 that the user's behavior risk score R(t) is high, so it enters step S3 to adjust the permissions.
[0057] During this process, the system will calculate the optimal permission adjustment strategy for the user through the optimal control algorithm; the system evaluates the current state x(t) and calculates the optimal control input u(t) to decide whether to reduce the user's access rights or require secondary identity authentication; after calculation by the optimal control algorithm, the system may tighten the user's access rights to sensitive resources and reduce the user's resource usage to ensure system security.
[0058] S4. Use game theory models to optimize the behavioral interaction between the system and users and solve the optimal security protection strategy. In step S4, the system optimizes the behavioral interaction between the system and users through game theory models to solve the optimal security protection strategy. This step follows the previous reinforcement learning model and optimal control theory, and relies on the framework of game theory to further analyze the game relationship between the system and users. The core goal of this process is to solve the optimal protection strategy of the system by calculating the interactive force between the system and users, so as to ensure the security of the system and improve the utilization efficiency of resources. The game theory model can provide an adaptive and optimized strategy solution by considering the interaction between the system's defense behavior and the user's resource access needs.
[0059] In this embodiment, step S4 models the behavior between the system and the user through a game theory model; specifically, the system and the user are regarded as participants in the game, the goal of the system is to maximize security, and the goal of the user is to maximize access to resources; game theory provides a framework for the system, in which the decisions of the system and the user influence each other through the game, thereby solving the optimal security protection strategy.
[0060] During the implementation process, the system inputs the results of risk assessment of user behavior (step S1) and permission adjustment (step S2), and solves the optimal strategy for the interaction between the system and the user through the framework of game theory; specifically, the system will use the game theory model to solve the optimal security protection measures based on the user's risk score R(t) and permission adjustment results, so as to formulate a strategy that does not affect the user experience while ensuring security.
[0061] The core content of the game theory model is the strategic interaction between the system and the user, and the security protection strategy is optimized by solving the Nash equilibrium. Nash equilibrium is a concept in game theory, which means that in a multi-party game, the strategy of each participant is optimized and cannot obtain higher utility by changing its own strategy. In other words, Nash equilibrium ensures the optimal interaction between the system and the user.
[0062] In this embodiment, the game between the system and the user can be represented by a utility function, which measures the "benefits" of the system and the user under a certain strategy; specifically: System utility function: The goal of the system is to maximize its own security while limiting user access when necessary; User utility function: The user's goal is to maximize access to system resources and bypass restrictions as much as possible; The utility function U of the system sys (t) is used to measure the security and defense effect of the system under a certain strategy; in this embodiment, the system utility function takes into account the security score of the system and the degree of control over user rights; specifically, the utility function of the system can be expressed as: U sys (t) = S(t) - λ·C(t); Where: S(t) represents the security score of the system at time t. The higher it is, the more secure the system is. C(t) represents the degree of control over user permissions by the system. The higher it is, the greater the degree of permission reduction is. λ is the weight coefficient, which controls the trade-off between security and permission control. Generally, the larger the λ is, the more the system will focus on improving security. The smaller the λ is, the more the system will focus on user experience.
[0063] The user's utility function U sys (t) measures the user's resource access under the current policy; the user's goal is to maximize the access rights to resources while minimizing the restrictions controlled by the system; specifically, the user's utility function is expressed as: U sys (t) = α·R(t)-β·P(t); Where: R(t) represents the risk score of user behavior, the higher the score, the greater the risk of user behavior; P(t) represents the degree of permission adjustment of the user at time t, the larger the score, the more resources the user can access; α and β are weight coefficients, which represent the user's sensitivity to behavior risk and permission adjustment, respectively.
[0064] In this embodiment, the user's goal is to maximize his access to resources R(t) and try to avoid the inconvenience caused by permission adjustment (quantified by P(t)); the game between the system and the user is balanced through the interaction of the user utility function and the system utility function.
[0065] Through the game theory model, the game between the system and the user finds the optimal strategy by solving the Nash equilibrium; Nash equilibrium plays a vital role in the game, which means that both the system and the user make the best choice based on the other party's strategy, and neither party can get a better result by unilaterally changing its own strategy; In this embodiment, the game between the system and the user can be expressed in the following form: maximize U sys (t), subject to U sys (t)≥U sys (t ′ ); The goal of this optimization problem is to maximize the system's utility function U sys (t), and at the same time satisfy a constraint: the system utility at the current time t is greater than or equal to that at an earlier time t ′ The utility of maximize U user (t), subject to U user (t)≥U user (t ′ ); The goal of this problem is to maximize the user's utility function U user (t), there is also a constraint: the user utility at the current time t is greater than or equal to that at some earlier time t ′ The utility of Among them, t ′ For possible strategy selection; the strategy selection of the system and the user will be based on the principle of maximizing their respective utility, and the Nash equilibrium point will be solved through game theory to determine the optimal security protection strategy.
[0066] After solving the Nash equilibrium of game theory, the system outputs the most suitable security protection measures for the current environment according to the optimal strategy; these measures include but are not limited to: Reduce user permissions: When the system detects abnormal user behavior or a high risk score, the system may limit the user's access to sensitive resources; Extending user rights: When the system assesses that the user's behavior is normal and the risk is low, the system will appropriately expand the user's access rights and optimize resource usage; Maintain existing permissions: For users with normal behavior and low security risk, the system will keep their current permissions unchanged.
[0067] These optimization measures ensure that the system can maintain high security while effectively utilizing resources and avoiding resource waste caused by excessive permission reduction.
[0068] Assume that there is a game between the system and user A, and user A frequently accesses highly sensitive resources and logs in through an abnormal IP at a certain moment; after evaluation, the system finds that user A's behavior risk is high, thus triggering permission reduction; in step S4, the system uses the game theory model to play a game with user A and calculates the strategy for user A to maximize his utility in this situation; finally, the system decides to reduce user A's access rights to sensitive resources and strengthen monitoring based on the game results.
[0069] In this process, the game theory model ensures that the system meets the user's resource needs to the greatest extent while providing protection, thus achieving a balance between security and user experience.
[0070] S5. Through the adaptive feedback mechanism, according to the real-time feedback data of the system operation, adjust the key parameters in the reinforcement learning model, optimal control model and game theory model to improve the responsiveness and flexibility of the system; In step S5, the system dynamically adjusts the model parameters according to real-time feedback through an adaptive feedback mechanism, thereby improving the flexibility and response speed of the system; this step ensures that the system can self-optimize according to the changing security environment and user behavior, and quickly respond to new security threats and resource requirements to maintain efficient operation of the system; in the aforementioned steps S1, S2 and S3, the system has achieved basic permission adjustment and security optimization through risk assessment, reinforcement learning and optimal control theory; step S5 further continuously optimizes the model through a feedback mechanism to ensure that the system can dynamically adapt to various environmental changes.
[0071] In this embodiment, step S5 realizes dynamic adjustment of model parameters through an adaptive feedback mechanism; specifically, the system adjusts the parameters in the reinforcement learning, optimal control and game theory models according to real-time user behavior data, risk assessment results and system performance feedback; the adjustment of these parameters helps the system improve its ability to respond to new threats, while optimizing the use of resources and avoiding excessive restrictions on user access rights.
[0072] During the implementation process, the system dynamically adjusts the following parameters based on the current system status and user behavior through real-time monitoring and feedback mechanisms: Risk threshold: Adjust the risk threshold R based on real-time user behavior data and security assessment. th , which affects the contraction and expansion of user permissions; for example, when user behavior frequently shows abnormalities, the system will automatically adjust the risk threshold and adopt stricter permission control.
[0073] Optimal control weight coefficient w 1 and w 2:The weight coefficients in the optimal control model control the balance between safety and resource utilization; the system adjusts these coefficients through real-time feedback to ensure that the system can maintain optimal safety and resource allocation in various scenarios.
[0074] Reinforcement learning reward function coefficients α and β: The reward function coefficients in the reinforcement learning model affect the system's penalty strategy for abnormal user behavior and access rights; based on feedback results, these coefficients are dynamically adjusted to optimize the learning strategy so that the system can adapt to new user behavior patterns.
[0075] These adjustments ensure that the system can flexibly respond to changing security threats and maximize resource utilization efficiency.
[0076] The working principle of the adaptive feedback mechanism is based on the collection, evaluation and optimization of real-time data; whenever the system detects a change in user behavior or system status, the feedback mechanism will go through the following steps: Real-time calculation of feedback values: By real-time monitoring of user behavior (such as login frequency, number of visits to sensitive resources, resource usage, etc.) and the system's security score, feedback values ΔF(t) and ΔS(t) are calculated to reflect the difference between the current strategy and the expected effect.
[0077] Adjust risk threshold: Based on real-time feedback data, the system dynamically adjusts the risk threshold R th , allowing the system to flexibly respond to different risk levels; for example, if the user's behavior pattern changes significantly, the system will adjust the risk threshold based on the changes to ensure the timeliness and accuracy of permission adjustments.
[0078] Optimize model parameters: The system adjusts the parameters in the optimal control and reinforcement learning models based on the feedback value; specifically, by adjusting α and β in real time, the system can update the abnormality of user behavior based on the latest behavior data and optimize the reward function; adjust the weight coefficient w 1 and w 2 Ensure the optimal balance between security and resource usage.
[0079] In order to make the feedback mechanism more accurate, the system dynamically updates key parameters through formulas; the following are the main formulas and parameter definitions involved in the dynamic update process: Dynamically update risk thresholds: ΔR th (t) = λ 1 ·(R(t)-R target (t)); Among them, ΔR th (t) represents the change at time t; R(t) is the risk score calculated in real time, R target (t) is the target risk score, λ1 It is the adjustment coefficient; the system controls the contraction and expansion of user permissions by adjusting the risk threshold.
[0080] Dynamically update the optimal control coefficient: Δw 1 (t) = λ 2 ·(S(t)-S target (t)); Δw 2 (t) = λ 3 ·(P(t)-P target (t)); Among them, Δw 1 (t) and Δw 2 (t) is the change at time t, S(t) is the current security score of the system, S target (t) is the target security score, P(t) is the current permission adjustment degree, and P target (t) is the target authority adjustment value, λ 2 and λ 3 are adjustment coefficients; by adjusting these coefficients, the system optimizes the balance between security and resource consumption; Dynamically adjust the coefficients of the reinforcement learning reward function: Δα(t)=λ 4 ·(R(t)-R target (t)); Δβ(t)=λ 5 ·(P(t)-P target (t)); Among them, Δα(t) and Δβ(t) are the changes at time t, R(t) and P(t) represent the current behavior abnormality score and authority adjustment degree respectively, R target (t) and P target (t) is the target score, λ 4 and λ 5 are adjustment coefficients; these adjustments ensure that the system can accurately adjust the learning strategy in the face of changes in user behavior and resource access.
[0081] In one possible implementation, the system can further improve the effectiveness of the adaptive feedback mechanism by introducing a reinforcement learning mechanism; reinforcement learning is a learning method based on exploration and feedback, in which the system continuously explores different control strategies and adjusts them based on each feedback.
[0082] Specifically, the system will adjust the parameters of the reinforcement learning model based on real-time feedback data (such as behavior changes, access to resources, etc.), so that the system can more accurately evaluate user behavior and adjust permissions in a timely manner; through reinforcement learning, the system can quickly adapt to new security threats in a changing environment and avoid vulnerabilities caused by static decisions.
[0083] Suppose the system is monitoring a user who usually accesses low-sensitivity resources, but suddenly accesses a large amount of sensitive data during a certain operation and attempts to bypass the access control policy; the system evaluates the user's behavior risk through step S1 and finds that the risk score exceeds the set threshold R th ; In step S2, the reinforcement learning model dynamically shrinks the user's access rights and starts the feedback mechanism; in step S5, the system adjusts the optimal control coefficient w according to the real-time behavior data and security assessment. 1 and w 2 , strengthen the balance between security and permission adjustment, and further optimize the permission adjustment strategy through reinforcement learning mechanism; this real-time feedback mechanism ensures that the system can respond to new security threats in a timely manner and optimize resource allocation.
[0084] See also Figure 2 , system safety management system, including: Behavior monitoring module, the behavior monitoring module is a key component of the entire system, used to collect and monitor user behavior feature data in real time; this module records and analyzes each user's behavior by tracking the user's operation behavior (such as login information, access frequency, operation type, resource access mode, etc.); by inputting these data into the behavior analysis model, the module can identify normal and abnormal user behavior and evaluate its potential security risks; when the system faces abnormal behavior, the module will provide real-time feedback and evaluation results to provide decision support for subsequent permission adjustment and security protection; permission adjustment module, the permission adjustment module dynamically adjusts the user's access rights based on reinforcement learning and optimal control models; this module calculates the user's risk score by real-time monitoring and analysis of the user's behavior data, and then decides whether to shrink or expand the user's access rights; the reinforcement learning model adaptively adjusts the access policy by continuously learning historical behavior data and real-time risk assessment results; the optimal control model ensures that the system maintains a balance between security and resource utilization efficiency by optimizing the cost function; the goal of this module is to adjust permissions in a timely manner according to changes in user behavior to avoid excessive restrictions while ensuring system security; Game theory model module: The game theory model module solves the optimal security protection strategy by optimizing the behavioral interaction between the system and the user. In a multi-user environment, there is a complex game relationship between the system and the user: the system needs to ensure high security, while the user seeks to maximize resource access. The game theory model calculates the utility of the system and the user under different strategies by modeling this interactive relationship, and finds the optimal strategy by solving the Nash equilibrium. This module helps the system to effectively allocate resources while ensuring security, balance the system's security requirements and the user's access requirements, and ensure the optimal protection effect. Adaptive feedback module, the adaptive feedback module adjusts model parameters based on real-time feedback data to improve the flexibility and response speed of the system; this module dynamically adjusts key parameters in the reinforcement learning model, optimal control model and game theory model, such as reward function coefficients, weight coefficients and risk thresholds, by continuously receiving feedback information during system operation; the system optimizes access control policies in real time according to changes in feedback data to cope with environmental changes or new security threats; this module ensures that the system can respond flexibly and quickly adjust policies when facing different scenarios, thereby maximizing resource utilization efficiency while improving security.
[0085] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A system security management method, characterized in that: The following steps are involved: Monitor user behavior data in real time, assess its security risks, and provide assessment results; Based on the evaluation results, the user's access rights are dynamically adjusted through the reinforcement learning model; Based on optimal control theory, optimize the user's access rights adjustment strategy to maximize the system's security and resource utilization; Use game theory models to optimize the behavioral interaction between the system and users and find the optimal security protection strategy; Through the adaptive feedback mechanism, key parameters in the reinforcement learning model, optimal control model and game theory model are adjusted according to the real-time feedback data of the system operation to improve the responsiveness and flexibility of the system.
2. The system security management method according to claim 1, characterized in that: The step of real-time monitoring of user behavior data includes: Obtain the user's login information, operation frequency, and access resource type and behavior characteristics; Combine users’ historical behavior data with real-time behavior data to assess potential security risks; The user behavior data includes but is not limited to: User's login information; Operating frequency; The type of resource being accessed; User's historical behavior data; Other relevant behavioral characteristics; The user behavior data is protected through encryption, desensitization and other technologies to ensure the privacy and compliance of the data.
3. The system security management method according to claim 1, characterized in that: The reinforcement learning model adjusts access permissions based on user behavior characteristics and real-time access requests by modeling the state space and action space of user behavior.
4. The system security management method according to claim 1, characterized in that: The reinforcement learning model uses the Q-learning algorithm to calculate the expected reward for each user behavior and adjusts user access rights in real time based on the reward.
5. The system security management method according to claim 1, characterized in that: The optimal control theory optimizes the access permission adjustment strategy by minimizing the cost function to balance system security and resource utilization. The cost function includes: Where: J(t) is the total cost function; S(t) is the security score of the system at time t; P(t) is the degree of adjustment of the system's permissions to the user; w1 and w2 are weight coefficients.
6. The system security management method according to claim 1, characterized in that: The steps of optimizing the user rights adjustment strategy using the optimal control theory include: Set the system state equation to represent the dynamic changes of user behavior and access requests; Real-time adjustment of user access rights is achieved by calculating the optimal control input.
7. The system security management method according to claim 1, characterized in that: The game theory model solves the Nash equilibrium and optimizes the protection and resource access game between the system and the user, so that the behavior between the system and the user is balanced.
8. The system security management method according to claim 1, characterized in that: The security protection strategies optimized by the game theory model include: The system maximizes security and users maximize access to resources; The strategies obtained through game analysis ensure security while not placing excessive restrictions on the user's operating experience.
9. The system security management method according to claim 1, characterized in that: The adaptive feedback mechanism includes: Adjust parameters in reinforcement learning, optimal control, and game theory models based on real-time risk assessment and behavioral data; Combine immediate feedback with historical behavior analysis to optimize user permission adjustment strategies.
10. A system security management system, applied to the system security management method according to any one of claims 1 to 9, characterized in that: include: Behavior monitoring module, used to monitor user behavior characteristic data in real time and conduct security risk assessment; The permission adjustment module is used to dynamically adjust user access rights based on reinforcement learning and optimal control models; Game theory model module, used to optimize the interaction between the system and users and solve the optimal protection strategy; Adaptive feedback module, which is used to adjust model parameters according to real-time feedback to improve the responsiveness and flexibility of the system.
Citation Information
Patent Citations
Network spoofing defense strategy optimization method and system based on intelligent real-time game
CN117220995A
Non-zero and multi-player game Q-learning method based on Gaussian process prediction
CN118732502A
Game strategy optimization method based on reinforcement learning
CN118940819A
Intelligent control method and system for switch cabinet
CN119002380A
Industrial internet of things embedded edge computer system
CN119416216A