Industrial internet cheating defense technology system based on supergame

Through the hypergame model and partially observable Markov decision model combined with the honeypot node and edge computing layer defense middleware, the defense strategy is optimized, and the problem of slow adjustment of defense strategy and insufficient resources in the existing technology is solved, and efficient and adaptive industrial Internet defense is achieved.

CN120498742APending Publication Date: 2025-08-15HUZHOU UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510594230.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing industrial Internet defense fraud technology is difficult to quickly adjust strategies, computing resources and real-time in complex environments, and there are insufficient adaptability and configuration complexity, which affects the overall defense effect.

Method used

The hypergame model and partially observable Markov decision model are adopted, combined with the defense middleware of honeypot nodes and edge computing layer, interfere with attackers through virtual data interaction and optimize defense strategies, and adopt zero-trust mechanism and deep reinforcement learning for adaptive defense.

Benefits of technology

It improves the security protection capabilities of the industrial Internet environment, optimizes defense strategies, improves decision-making efficiency, achieves cost balance, and can effectively respond to a variety of advanced and lasting threats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498742A_ABST
    Figure CN120498742A_ABST
Patent Text Reader

Abstract

The invention discloses a supergame-based industrial internet cheating defense technology system, which comprises an industrial control layer which is a distributed network comprising a plurality of nodes; the nodes are industrial internet equipment nodes or honeypot nodes; the honeypot node is used for generating a bait environment for virtual data interaction; the multi-access edge computing layer comprises a plurality of edge computing nodes, and the edge computing nodes deploy deception prevention middleware and are used for carrying out security authentication on the data access request; and after a data access request proposed by the nodes of the industrial control layer is subjected to security authentication through the nodes of the multi-access edge computing layer, data services are provided for the nodes of the industrial control layer through the multi-access edge computing nodes. According to the method, the supergame theory and the partially observable Markov decision model are combined, the security protection capability in the industrial internet environment is effectively improved, and the problems of dynamically adapting to complex attacks, optimizing defense strategies, improving decision efficiency, realizing cost balance and the like in the industrial internet environment are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Internet of Things security technology, and more specifically, relates to an industrial Internet deception defense technology system based on hyper-game. Background Art

[0002] Faced with the increasingly complex security landscape of the Industrial Internet, research and application of related technologies are accelerating. Currently, the Industrial Internet security protection system is transitioning towards end-to-end comprehensive protection. The adoption of zero-trust architecture, multi-factor authentication, data encryption, and edge computing can significantly enhance security. With the deep integration of 5G and edge computing, the security needs of the Industrial Internet are becoming increasingly urgent, and anti-spoofing technologies have become a research hotspot. These innovative security technologies not only effectively address traditional cybersecurity threats but also provide countermeasures for complex threats in future IoT environments. The Industrial Internet applies IoT technology to industrial production, enabling autonomous sensing and intelligent collaboration in devices, broadening the application areas of industrial automation and intelligent manufacturing. Therefore, there is an urgent need to strengthen comprehensive security protection measures to address evolving security challenges.

[0003] Deception defense technology can effectively combat increasingly sophisticated cyberattacks within the Industrial Internet, particularly malicious attacks targeting critical infrastructure. Through deception defense, systems can introduce false or misleading information to confuse and mislead attackers, preventing them from directly accessing core system components or sensitive data, thereby protecting real resources. Anonymous deception defense technology within the Industrial Internet primarily addresses the potential risks associated with the exposure of decoy information during the deception defense process. It emphasizes the anonymity of decoy information, making it difficult for attackers to distinguish between fake and real system targets. At the same time, these decoys must possess sufficient authenticity to divert attackers away from the real system, slowing the attack and allowing defense systems to respond quickly, effectively protecting critical equipment and data resources within the Industrial Internet.

[0004] Currently, various institutions have disclosed methods for defending against deception technologies in the industrial internet. Patent application CN118821937A proposes a hyper-game-based computing network attack and defense model. However, this invention faces challenges such as high computational complexity, reliance on incomplete information for decision-making accuracy, insufficient adaptability, and model distortion caused by network simplification. Patent application CN110430190A proposes a deception defense system based on the ATT&CK framework. This system aims to integrate deception defense technologies with network assets to build a comprehensive defense mechanism. However, implementation may face challenges such as database construction complexity, environmental adaptability challenges, and reliance on attacker behavior prediction. Patent application CN114465747A discloses an active deception defense method and system based on dynamic port masquerading. By deploying honeypots and dynamic port configuration, this method effectively confuses attackers and protects real nodes. However, this method may face challenges such as configuration complexity, service response delays, and adaptability to dynamic environments, which may affect the overall defense effectiveness. Patent application CN118827216A discloses an active deception defense method based on dynamic port masquerading. By deploying honeypots and dynamic configuration, it effectively confuses attackers and protects real nodes, but may face challenges in configuration complexity and environmental adaptability. Patent application CN118827216A discloses a spacecraft model predictive control method that uses a robust algorithm to effectively defend against deception attacks and ensure normal operation, but may face challenges such as model complexity and dynamic adaptability.

[0005] In summary, existing industrial internet deception defense technologies generally face numerous challenges in practical application. On the one hand, various defense strategies must be effectively implemented in complex environments, and rapidly selecting and adjusting strategies in different attack scenarios remains a significant challenge. On the other hand, many defense mechanisms face limitations in computing resources and real-time performance. Especially when employing complex algorithms, balancing defense effectiveness with computational cost is a pressing issue. Furthermore, existing technologies also exhibit shortcomings in adaptability, configuration complexity, and responsiveness to environmental changes, impacting overall defense effectiveness. Summary of the Invention

[0006] In response to the above defects or improvement needs of the prior art, the present invention provides an industrial Internet defense deception technology system based on hyper-game, which aims to use a hyper-game model to simulate complex and dynamically adjusted long-term advanced attacks (advanced persistent threat APT), use honeypot nodes deployed in the industrial control layer to induce interaction, and use the defense middleware deployed in the multi-access edge computing layer to make defense strategy decisions based on a partially observable Markov decision model, so as to adapt to the complex and long-term dynamically adjusted security threat environment and optimize the defense effect, thereby solving the technical problem that the existing industrial Internet of Things defense system has poor defense effect under long-term multi-stage attacks.

[0007] To achieve the above objectives, according to one aspect of the present invention, there is provided an industrial Internet deception defense technology system based on hyper-game, which includes an industrial control layer, a multi-access edge computing layer, and a backend data layer;

[0008] The industrial control layer is a distributed network including multiple nodes; the nodes are industrial Internet device nodes or honeypot nodes; the honeypot nodes are used to interfere with the attacker's identification of the real target and capture the characteristics of the attack behavior by creating a bait environment for virtual data interaction;

[0009] The multi-access edge computing layer includes a plurality of edge computing nodes, wherein the edge computing nodes deploy anti-spoofing middleware for performing security authentication on data access requests;

[0010] The backend data layer is used to store data required for the operation of the Industrial Internet and provide data services based on data access requests;

[0011] The data access request raised by the node of the industrial control layer is securely authenticated by the node of the multi-access edge computing layer, and then the multi-access edge computing node provides data services to the node of the industrial control layer.

[0012] Preferably, in the hyper-game-based industrial Internet deception defense technology system, the edge computing node adopts a zero-trust mechanism for access to sensitive data. When the edge computing node receives a data access request, it formulates a current defense strategy and executes a defense action.

[0013] Preferably, the industrial Internet defense deception technology system based on hyper-game, its defense middleware perceives the current attack stage of all attackers x in the hyper-game chain based on multi-source data and evaluates the uncertainty of the perception results, adopts a decision model based on a partially observable Markov decision process to formulate the optimal defense strategy of the current hyper-game chain, and executes defense actions.

[0014] Preferably, the hyper-game-based industrial Internet deception defense technology system includes the following five attack phases according to MITRE's ATT&CK framework:

[0015] Initial access, which refers to the stage when an attacker first breaches the target network perimeter and establishes a foothold;

[0016] Defense evasion refers to techniques used by attackers to conceal malicious behavior and bypass security detection;

[0017] Credential access, which refers to obtaining user account passwords, tokens, or hashes to escalate privileges;

[0018] Lateral movement, where an attacker spreads control within a target network and gains access to core data;

[0019] Command control, where the attacker communicates with the controlled host through the server and issues instructions;

[0020] Treat each attack phase as a subgame of the supergame.

[0021] Preferably, the industrial Internet deception defense technology system based on hyper-game, wherein the defense middleware perceives the current attack stage of the hyper-game chain based on multi-source data, specifically:

[0022] The defense middleware uses a machine learning classification algorithm or a rule engine classification to determine the current attack stage based on the detection results of the intrusion detection system, log analysis, and / or threat intelligence; wherein the intrusion detection system is used to detect known attack patterns or abnormal traffic; the log analysis specifically monitors system logs and network traffic data to identify abnormal behavior; the threat intelligence is attack behavior information updated by an external threat intelligence source.

[0023] Preferably, the industrial Internet defense deception technology system based on hypergame, wherein the defense middleware evaluates the uncertainty of the perception result, specifically calculates the uncertainty of the perception attacker x in the subgame h in the following way

[0024]

[0025] Where ζ is a positive constant used to adjust the size and impact of the attacker's exposure to communication performance. Dx represents the degree of leakage of the defense deception strategy detected by the defense middleware. During the interaction process, the defense middleware simulates the attacker observing abnormal changes in system behavior and uses the ratio of the current environmental characteristic value to the expected environmental characteristic value without the defense deception strategy to represent the degree of leakage of the defense deception strategy. The environmental characteristic value is the weighted sum of the values of various environmental characteristics, including response time, data packet characteristics, and return information success rate. Attackers can exploit the leakage of the defense deception strategy and adjust their own attack strategy to deal with this defense mechanism, thereby introducing uncertainty in the defense middleware's perception. It represents the length of time the defense middleware observes the attacker. The longer the observation time, the smaller the perceived uncertainty. f(u) represents the impact of communication performance on the attacker. When the f(u) value is smaller, the communication stability is higher. It is defined as follows:

[0026] f(u)=(1+BW n +DL+PL) / 3

[0027] Among them, BW n It is the bandwidth utilization rate of communication between the edge computing nodes of the multi-access edge computing layer and the industrial Internet device node n of the currently communicating industrial control layer; DL is the total transmission delay between the edge computing nodes of the multi-access edge computing layer and the industrial Internet device node n of the currently communicating industrial control layer; PL is the packet loss rate between the edge computing nodes of the multi-access edge computing layer and the industrial Internet device node n of the currently communicating industrial control layer.

[0028] Preferably, the industrial Internet deception defense technology system based on hyper-game adopts a decision model based on a partially observable Markov decision process to formulate the optimal defense strategy of the current hyper-game chain, specifically:

[0029] The defense middleware will calculate the attack stage h of all current attackers x and the perceived uncertainty As the current state, formulate the current defense strategy for node n based on the corresponding subgame h of the super game, and select the total expected defense utility value of the super game Minimal defensive strategy and execution of defensive moves;

[0030] The total expected defense utility value of the super game is It can be evaluated by adding the attack utilities of all attackers in the Nash equilibrium state of subgame h.

[0031] Preferably, the hyper-game-based industrial Internet deception defense technology system, wherein the decision model based on the partially observable Markov decision process is defined as follows:

[0032] State S, including the communication performance indicators of the Industrial Internet, the perceived attack stage h, and the uncertainty of the attacker x in the defensive middleware perception subgame h The attack uncertainty of attacker x in a given subgame h Current state of defensive deception technology enabled d, system beliefs;

[0033] Action space A, i.e., the set of defense strategies;

[0034] Reward function r t , based on the expected utility, the Bellman equation is used to discount the future expected value of the reward and perform long-term optimal strategy planning, which is recorded as:

[0035]

[0036] Here, γ is a discount factor that is used to balance short-term and long-term rewards.

[0037] Preferably, the communication performance indicators of the industrial Internet anti-deception technology system based on hyper-game include the data transmission rate R of each link. n (t), bandwidth utilization BW n , total transmission delay DL, packet loss rate PL;

[0038] System belief, which represents the true probability of adopting defense strategy t in the Industrial Internet neutron game h. Calculate as follows:

[0039]

[0040] in, The probability that the defensive intermediate adopts the defensive strategy t to respond to each node in the subgame h is, To defend the intermediate's initial belief about the environment and the opponent's behavior before the game begins, It refers to the total probability of responding by considering all possible defense strategies t;

[0041] The attack uncertainty of attacker x in a given subgame h Calculate as follows:

[0042]

[0043] Among them, d is the defensive deception technology enabled status, which represents whether the system is using defensive deception technology to deal with potential attacks. When d is set to 1, it means that the system is actively using defensive deception technology to improve the speed and accuracy of response to attacks; when d is 0, it means that the system has not enabled this technology; It refers to the time that attackers monitor industrial control systems. The acquisition of monitoring time requires a combination of passive detection and active induction for perception.

[0044] Preferably, the industrial Internet deception defense technology system based on hyper-game adopts the CG-D3QN network based on the D3QN network as the decision model based on the partially observable Markov decision process; the CG-D3QN network includes: state strategy network parameters The value function network θ1 is used to calculate the Q value and combine the loss function to back-propagate and train. The gradient descent method is often used to update the loss function. The value function network θ1 is used to predict the current Q value of the defense middleware. The target network θ2 is used to predict the Q value of the defense middleware at the next moment. The action advantage network Δ is used to judge the value difference of different actions under the current value function network θ1. And the state value network ψ is used to evaluate the current state S t The action advantage network Δ and the state value network ψ are jointly updated according to the Q-value error using back propagation combined with gradient descent.

[0045] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0046] The defensive deception technology system provided by the present invention combines hyper-game theory and partially observable Markov decision models, effectively improving the security protection capabilities in the industrial Internet environment, and solving problems such as dynamically adapting to complex attacks, optimizing defense strategies, improving decision-making efficiency, and achieving cost balance in the industrial Internet environment.

[0047] The method provided by this invention continuously optimizes the system's decision-making capabilities during defense strategy training through an adaptive learning mechanism, addressing the rigidity and inefficiency of traditional defense models. This method continuously enhances the flexibility of defense behaviors through reinforcement learning, effectively addressing a wide range of advanced persistent threat attacks.

[0048] The optimal technical solution, by combining hyper-game and deep reinforcement learning, enables the system to adjust defense strategies in real time and automatically update defense parameters according to different attack modes, greatly improving the adaptive ability and long-term stability of the defense system. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a schematic diagram of the structure of the industrial Internet anti-deception technology system based on hyper-game provided by the present invention;

[0050] Figure 2 This is a graph showing the mean time between failures test results of the industrial Internet anti-deception technology system provided by an embodiment of the present invention;

[0051] Figure 3This is a graph showing the false alarm rate test results of the industrial Internet anti-deception technology system provided by an embodiment of the present invention;

[0052] Figure 4 It is the expected attack utility of the industrial Internet defense deception technology system provided by the embodiment of the present invention;

[0053] Figure 5 This is the expected defense utility of the industrial Internet anti-deception technology system provided by the embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the following embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0055] The present invention provides an industrial Internet deception defense technology system based on hyper-game, such as Figure 1 As shown, it includes the industrial control layer, the multi-access edge computing layer, and the backend data layer;

[0056] The industrial control layer is a distributed network including multiple nodes; the nodes are industrial Internet device nodes or honeypot nodes; the honeypot nodes are used to interfere with the attacker's identification of the real target and capture the characteristics of the attack behavior by creating a bait environment for virtual data interaction;

[0057] The multi-access edge computing layer includes a plurality of edge computing nodes, wherein the edge computing nodes deploy anti-spoofing middleware for performing security authentication on data access requests;

[0058] The backend data layer is used to store data required for the operation of the Industrial Internet and provide data services based on data access requests;

[0059] The data access request proposed by the node of the industrial control layer is securely authenticated by the node of the multi-access edge computing layer, and the multi-access edge computing node provides data services to the industrial control layer node;

[0060] The defense middleware perceives the current attack stage of all attackers x in the hyper-game chain based on multi-source data and evaluates the uncertainty of the perception results. It adopts a decision model based on a partially observable Markov decision process to formulate the optimal defense strategy for the current hyper-game chain and executes defense actions. Preferably, the edge computing node adopts a zero-trust mechanism for access to sensitive data. When the edge computing node receives a data access request, it formulates the current defense strategy and executes a defense action.

[0061] According to MITRE's ATT&CK framework, the attack phase includes the following five:

[0062] Initial access (IA), which refers to the stage when an attacker first breaks through the target network perimeter and establishes a foothold;

[0063] Defense evasion (DE) refers to techniques used by attackers to conceal malicious behavior and bypass security detection;

[0064] Credential access (CA), which refers to obtaining user account passwords, tokens, or hashes to elevate privileges;

[0065] Lateral movement (LM) refers to an attacker spreading control within the target network and gaining access to core data;

[0066] Command and control (C2) refers to the attacker communicating with the controlled host through the server and issuing instructions;

[0067] Treating each attack phase as a subgame of the hypergame reflects the complexity of IIoT APT strategies, where attackers perform multiple actions to achieve their goals. Attackers include both internal and external attackers.

[0068] The defense middleware perceives the current attack stage of the hyper game chain based on multi-source data, specifically:

[0069] The defense middleware uses a machine learning classification algorithm or a rule engine classification to determine the current attack stage based on the detection results of the intrusion detection system, log analysis, and / or threat intelligence; wherein the intrusion detection system (IDS) is used to detect known attack patterns or abnormal traffic; the log analysis specifically monitors system logs and network traffic data to identify abnormal behavior; the threat intelligence is attack behavior information updated by an external threat intelligence source.

[0070] The defense middleware evaluates the uncertainty of the perception result, specifically calculating the uncertainty of the perceived attacker x in the subgame h as follows

[0071]

[0072] Where ζ is a constant, a positive number, used to adjust the size and impact of the attacker's exposure to communication performance. Dx represents the degree of leakage of the defense deception strategy detected by the defense middleware. During the interaction process, the defense middleware simulates the attacker observing abnormal changes in system behavior and uses the ratio of the current environmental characteristic value to the expected environmental characteristic value without the defense deception strategy to represent the degree of leakage of the defense deception strategy. The environmental characteristic value is the weighted sum of the values of various environmental characteristics, including response time, data packet characteristics, and return information success rate. Attackers can exploit the leakage of the defense deception strategy and adjust their own attack strategy to deal with this defense mechanism, thereby introducing uncertainty in the defense middleware's perception. It represents the length of time the defense middleware observes the attacker. The longer the observation time, the smaller the perceived uncertainty. f(u) represents the impact of communication performance on the attacker. When the f(u) value is smaller, the communication stability is higher. It is defined as follows:

[0073] f(u)=(1+BW n +DL+PL) / 3

[0074] Among them, BW n is the bandwidth utilization of communication between the edge computing node of the multi-access edge computing layer and the industrial Internet device node n of the currently communicating industrial control layer; DL is the total transmission delay between the edge computing node of the multi-access edge computing layer and the industrial Internet device node n of the currently communicating industrial control layer; PL is the packet loss rate between the edge computing node of the multi-access edge computing layer and the industrial Internet device node n of the currently communicating industrial control layer;

[0075] Bandwidth utilization BW of industrial Internet device node n n , the calculation method is as follows:

[0076]

[0077] Where T is the total period, R n (t) represents the transmission rate of the industrial Internet device node n in the industrial control layer in time slot t, R(t) is the transmission rate of the industrial Internet device node in the industrial control layer in time slot t, and R(t) is the uplink transmission rate R u (t) and the downlink transmission rate R d The sum of (t) is calculated as follows:

[0078]

[0079] Where B is the channel bandwidth, P u and P d are the transmission power of industrial Internet device nodes and edge computing nodes, h u and hd are the channel gains of the uplink channel and the downlink channel respectively, and N0 is the noise power spectral density;

[0080] The total transmission delay DL is calculated as follows:

[0081] DL=D trans +D prop +D proc

[0082] DL refers to the total time it takes for data to travel from the sender to the receiver, and consists of the following three main parts: trans is the transmission delay, D prop is the propagation delay, D proc To process delays.

[0083] The packet loss rate PL is calculated as follows:

[0084]

[0085] Among them, P lost is the number of lost packets, P total The total number of packets transmitted.

[0086] The decision model based on partially observable Markov decision process is used to formulate the optimal defense strategy of the current hypergame chain, specifically:

[0087] The defense middleware will calculate the attack stage h of all current attackers x and the perceived uncertainty As the current state, formulate the current defense strategy for node n based on the corresponding subgame h of the super game, and select the total expected defense utility value of the super game Minimal defensive strategy and perform defensive moves.

[0088] The total expected defense utility value of the super game is It can be evaluated by adding the attack utilities of all attackers in the Nash equilibrium state of subgame h. The specific calculation method is as follows:

[0089]

[0090] in, is the expected attack utility of attacker x in the Nash equilibrium state of subgame h, which is calculated as follows:

[0091]

[0092] Among them, A sTo represent the attack strategy adopted by attacker x in the Nash equilibrium state of subgame h, the Nash equilibrium state can be solved according to the specific game type, or by simulating the attack and defense process to a stable state; δ is the discount factor, which ranges from [0,1]. It refers to the comprehensive utility function of attacker x in the defense middleware-aware subgame h, which is expressed as follows:

[0093]

[0094] in, is the strategy set C of attacker x against the defense middleware in subgame h I,h The expected utility on the defense strategy that the defense middleware can select from the strategy set according to the prior probability p is calculated as follows:

[0095]

[0096] Among them, p i is the probability of state i occurring; u i (i,SA,S D ) is the attacker in state i, taking strategy A s The defense strategy is S D The utility of time; s is the strategy chosen by the attacker, C I,h represents the possible strategies that the defense middleware may choose in subgame h.

[0097] is the expected utility of the attacker when the defense middleware fails to sense the failure and selects the strategy with the minimum defense utility for each node in a given subgame h. In this case, the defense middleware will choose the strategy with the minimum defense utility and the maximum attacker utility. Denoted as:

[0098]

[0099] The probability of the minimum expected utility strategy for defensive middleware is p i , and for these worst strategies, the attacker's utility is u i ,but:

[0100]

[0101] Among them, p i is the probability of node state i occurring when the defense middleware selects the minimum expected utility strategy; The attacker is in state i and adopts strategy S A , while the defensive middleware takes The utility of time; s Still the attacker's strategy of choice, CMSw,h represents the worst set of strategies selected by the defense middleware in subgame h.

[0102] First, for each possible sub-game stage h, the attacker will evaluate his specific strategy set C in the defense middleware m,n The expected utility under:

[0103]

[0104] The attacker has a fixed defense strategy C m,n Under this condition, choose the optimal attack strategy that maximizes its expected utility

[0105]

[0106] Defense middleware anticipates that attackers will choose the above Therefore, its goal is to select the optimal defense strategy C based on this m,n , to minimize the attacker's maximum expected utility or maximize one's own utility:

[0107]

[0108] or equivalently:

[0109]

[0110] Finally, the obtained strategy combination The following equilibrium conditions are met:

[0111]

[0112] Here, It is the Nash equilibrium point in this sub-game stage, and it is also the perfect equilibrium solution of the sub-game in the multi-stage game.

[0113] is the attack uncertainty of attacker x in a given subgame h. Assuming that the attacker is completely certain, then If you are not sure, set Calculate as follows:

[0114]

[0115] Among them, d is the defensive deception technology enabled status, which represents whether the system is using defensive deception technology to deal with potential attacks. When d is set to 1, it means that the system is actively using defensive deception technology to improve the speed and accuracy of response to attacks; when d is 0, it means that the system has not enabled this technology; It refers to the time an attacker monitors an industrial control system. The acquisition of monitoring time requires a combination of passive detection (logs, traffic) and active induction (deception technology) for perception.

[0116] The attacker's expected utility is the expected value of the attacker's expected utility. The attacker's attack utility is Calculate as follows:

[0117]

[0118] in, It's income, is the loss caused by the attack strategy s when the defense middleware adopts strategy t in node n. It is calculated as follows:

[0119]

[0120] Among them, the attack revenue Including attack impact ai based on the attacker's strategy s sn and the required defense cost dc in node n tn , the required defense cost dc in node n tn Determined by the defense strategy; attack loss includes based on the defense impact di tn and the attack cost ac sn , the attack cost ac sn Determined by the attack strategy detected by the intrusion detection system. tn and ac sn There are three levels: Low = 1; Medium = 2; High = 3;

[0121] The calculation method is as follows:

[0122]

[0123] Among them, A n represents the number of authentication attempts required for an attacker to compromise node n, ranging from [0,5]. Lsn represents the set of vulnerabilities in node n that can be exploited by an attacker. T ht represents the response time of the defense middleware’s strategy adjustment based on the complexity of the APT attack in subgame h, P hs refers to the probability of APT attackers, r refers to the normal number that adjusts the defense effect, ai tn represents the expected attack impact of node n when using defense strategy t. If no attacker is detected, then ai tn =0, ω refers to the weight of measuring the exploitability of the attack, I n The importance of node n is determined by the topology of the industrial control layer.

[0124] The attacker and the defensive middleman, as the two participants of the present invention, engage in a long-term and continuous game. There is uncertainty in the strategic choices of both parties. This uncertainty stems from the incomplete understanding of each other's predicted behavior. The probability that the defensive middleman knows that the attacker is in a certain attack stage (sub-game h) is in:

[0125] The decision model based on partially observable Markov decision process is defined as follows:

[0126] State S, including the communication performance indicators of the Industrial Internet, the perceived attack stage h, and the uncertainty of the attacker x in the defensive middleware perception subgame h The attack uncertainty of attacker x in a given subgame h Current state of defensive deception technology enabled d, system beliefs;

[0127] The communication performance indicators include the data transmission rate R of each link n (t), bandwidth utilization BW n , total transmission delay DL, packet loss rate PL;

[0128] System belief, which represents the true probability of adopting defense strategy t in the Industrial Internet neutron game h. Calculate as follows:

[0129]

[0130] in, The probability that the defensive intermediate adopts the defensive strategy t to respond to each node in the subgame h is, To defend the intermediate's initial belief about the environment and the opponent's behavior before the game begins, Refers to the total probability of responding by taking into account all possible defense strategies t; calculates the probability that the defense intermediate adopts the defense strategy s to respond to each node in the subgame h of the defense intermediate When the attacker successfully damages node n, the defense middleware perceives the uncertainty of the attacker x in the subgame h. Cannot be fully estimated Since the defense intermediate has its uncertainty about the attacker x It will be with probability Each node is observed correctly, and The probability of ignoring a node.

[0131] Action space A, i.e., the set of defense strategies;

[0132] Reward function r t, based on the expected utility, the Bellman equation is used to discount the future expected value of the reward and perform long-term optimal strategy planning, which is recorded as:

[0133]

[0134] Here, γ is a discount factor that is used to balance short-term and long-term rewards.

[0135] The preferred solution is to use a CG-D3QN network based on a D3QN network as the decision model based on the partially observable Markov decision process; the CG-D3QN network includes: state strategy network parameters The value function network θ1 is used to calculate the Q value and combine the loss function to back-propagate and train. The gradient descent method is often used to update the loss function. The value function network θ1 is used to predict the current Q value of the defense middleware. The target network θ2 is used to predict the Q value of the defense middleware at the next moment. The action advantage network Δ is used to judge the value difference of different actions under the current value function network θ1. And the state value network ψ is used to evaluate the current state S t The action advantage network Δ and the state value network ψ are jointly updated according to the Q-value error using back propagation combined with gradient descent.

[0136] For the CG-D3QN network, historical monitoring time data can be stored through the experience replay pool, and the network can be trained to learn attacker behavior patterns, thereby optimizing the timing of enabling defensive deception technology.

[0137] The following are examples:

[0138] The present invention provides an industrial Internet deception defense technology system based on hyper-game, such as Figure 1 As shown, it includes the industrial control layer, the multi-access edge computing layer, and the backend data layer;

[0139] The industrial control layer is a distributed network including multiple nodes; the nodes are industrial Internet device nodes or honeypot nodes; the honeypot nodes are used to interfere with the attacker's identification of the real target and capture the characteristics of the attack behavior by creating a bait environment for virtual data interaction;

[0140] The multi-access edge computing layer includes a plurality of edge computing nodes, wherein the edge computing nodes deploy anti-spoofing middleware for performing security authentication on data access requests;

[0141] The backend data layer is used to store data required for the operation of the Industrial Internet and provide data services based on data access requests;

[0142] The data access request proposed by the node of the industrial control layer is securely authenticated by the node of the multi-access edge computing layer, and the multi-access edge computing node provides data services to the industrial control layer node;

[0143] The defense middleware uses the CG-D3QN network decision-making defense strategy to explicitly model the interaction between attack and defense strategies, significantly improving the robustness of the defense middleware in dynamic confrontation. The specific steps are as follows:

[0144] (1) For each defense control middleware, observe and obtain the status s of the current time slot t , and estimate the uncertainty of the defense middleware perceiving the attacker x in the subgame h and the attack uncertainty of attacker x in a given subgame h

[0145] Status t Specifically including: communication performance indicators, attacker characteristics, current defense status, attack stage perceived by the defense middleware, and uncertainty about the attacker x in the defense middleware perception subgame h. The attack uncertainty of attacker x in a given subgame h The current state of defensive deception technology enabled, and system beliefs;

[0146] The communication performance indicators include the data transmission rate R of each link n (t), bandwidth utilization BW n , total transmission delay DL, and packet loss rate PL;

[0147] The attacker characteristics, including intrusion detection scores, abnormal traffic patterns, and attack event triggering frequencies, can be obtained using an intrusion detection system.

[0148] (2) Using D3QN to select a simulated attack strategy s for the attacker, and using the Q network θ1 to select the defense strategy t for the defense middleware response according to the simulated attack strategy s until the Nash equilibrium state is reached;

[0149] In the equilibrium state, the strategies of the attacker and the defensive middleware are mutually optimal responses, that is, neither party can improve its own utility by changing its strategy alone. Therefore, in the subgame Nash equilibrium, the equilibrium strategies of the attacker and the defensive middleware are Satisfies: Among them: Is the attacker's optimal strategy, that is, to take the defensive middleware, When , the attacker chooses the strategy that maximizes his own utility. It is the optimal strategy for defending middleware, that is, defending middleware when the attacker chooses When , the strategy that minimizes the attacker's utility. This equilibrium point This is the Nash equilibrium in the subgame. This equilibrium point represents the stable strategy choices of the attacker and the defensive middleware in the subgame, that is, when both parties know the other's strategy, they have no motivation to change their strategy.

[0150] To ensure the robustness and stability of defense strategies during multi-stage games, it is necessary to identify the optimal attack-defense strategy combination within a given subgame. This combination should satisfy the following conditions: neither player has an incentive to unilaterally deviate from their current strategy while the other player's strategy remains unchanged. Therefore, the equilibrium point of the game can be derived using the concept of a subgame-perfect Nash equilibrium.

[0151] The D3QN is a fully connected Dueling architecture, which uses experience replay technology (Replay Buffer) and uses the target network (Target Network) θ2 for update during training.

[0152] The defense strategy set in this embodiment i is a natural number from 1 to 8, to They represent firewall deployment, patching software vulnerabilities, updating encryption keys, expelling controlled nodes, hybrid honeypot deployment, information collection, injecting incorrect keys, and hiding network topology connections.

[0153] (3) Observe the network status and calculate the attack utility of the IIoT APT attacker and the defense utility of the corresponding defense middleware;

[0154] The attacker's attack utility Calculate as follows:

[0155]

[0156] in, It's income, is the loss caused by the attack strategy s when the defense middleware adopts strategy t in node n. It is calculated as follows:

[0157]

[0158] Among them, the attack revenue Including attack impact ai based on the attacker's strategy s sn and the required defense cost dc in node n tn , attack loss includes based on defense impact di tn and the attack cost ac sn . dc tn and ac snThere are three levels: Low = 1; Medium = 2; High = 3. The calculation method is as follows:

[0159]

[0160] Among them, A n represents the number of authentication attempts required for an attacker to compromise node n, ranging from [0,5]. Lsn represents the set of vulnerabilities in node n that can be exploited by an attacker. T ht represents the response time of the defense middleware’s strategy adjustment based on the complexity of the APT attack in subgame h, P hs refers to the probability of APT attackers, r refers to the normal number that adjusts the defense effect, ai tn represents the expected attack impact of node n when using defense strategy t. If there is no APT attacker or NIDS detects it, then ai tn =0, ω refers to the weight of measuring the exploitability of the attack, I n The importance of node n is determined by the topology of the industrial control layer.

[0161] Attack Type:

[0162] Software vulnerabilities (SV) refer to potential program errors or design flaws in the system. The SV of each node is modeled as a random variable that varies in the range of [0,1].

[0163] Key Vulnerabilities (EV) These vulnerabilities are caused by the leakage of encryption keys. As the time an internal attacker spends within the system increases, the likelihood that they can obtain encryption keys from other legitimate nodes also increases.

[0164] Other vulnerabilities (UV): Randomly change [0,10]

[0165] The attack strategy set in this embodiment Consider 8 attack strategies, to They are: vulnerability scanning, phishing attacks, botnet attacks, distributed denial of service attacks, zero-day vulnerability exploits, key exposure, privilege escalation, and data tampering; the attack strategies and their characteristics are shown in Table 1:

[0166] Table 1 Attack strategies and their characteristics

[0167]

[0168]

[0169] Attack cost: Low = 1; Medium = 2; High = 3.

[0170] (4) The state of the current time slot st , Defense strategy adopted by defense middleware a t , detect the attacker's attack strategy, current reward r t , and the state s of the next time slot t+1 As an experience, it is stored in the experience pool D for network parameter update; where:

[0171]

[0172] Here, γ is a discount factor that is used to balance short-term and long-term rewards.

[0173] The total expected defense utility value of the super game is It can be evaluated by adding the attack utilities of all attackers in the Nash equilibrium state of subgame h. The specific calculation method is as follows:

[0174]

[0175] in, is the expected attack utility of attacker x in the Nash equilibrium state of subgame h, which is calculated as follows:

[0176]

[0177] Among them, A S To represent the attack strategy adopted by attacker x in the Nash equilibrium state of subgame h, the Nash equilibrium state can be solved according to the specific game type, or by simulating the attack and defense process to a stable state; δ is the discount factor, which ranges from [0,1]. It refers to the comprehensive utility function of attacker x in the defense middleware-aware subgame h, which is expressed as follows:

[0178]

[0179] Among them, among them, is the strategy set C of attacker x against the defense middleware in subgame h I,h Based on the expected utility, the defense middleware can select the defense strategy in the strategy set according to the prior probability p; is the expected utility of the attacker when the defense middleware fails to sense the failure and selects the strategy with the minimum defense utility for each node in a given subgame h. In this case, the defense middleware will choose the strategy with the minimum defense utility and the maximum attacker utility. Denoted as:

[0180]

[0181] For any defense strategy set C m,n , The calculation method is as follows:

[0182]

[0183] Among them, p i is the probability of state i occurring; u i (i,S A ,S D ) is the attacker in state i, taking strategy A s And take S D The attack utility when A s is the strategy chosen by the attacker, C I,h represents the possible strategies that the defense middleware may choose in subgame h.

[0184] The defense strategy set in this embodiment Consider 8 defense strategies, to They are: firewall deployment, patching software vulnerabilities, updating encryption keys, expelling infected nodes, hybrid honeypot deployment, information collection, injecting false keys, and hiding network topology connections; the defense strategies and their characteristics are shown in Table 2:

[0185] Table 2: Key characteristics of defense strategies

[0186]

[0187] At different attack stages, the defense middleware formulates different defense strategies based on possible attack strategies. The following table specifically considers eight defense strategies.

[0188]

[0189] 8 defensive strategies to They are: firewall deployment, patching software vulnerabilities, updating encryption keys, expelling compromised nodes, hybrid honeypot deployment, information collection, injecting fake secret keys, and hiding network topology connections;

[0190] Update the Q network θ1 as follows:

[0191]

[0192] in, is the sample s in the experience pool t ,a t ,r t ,s t+1 The expectation of the square of the gradient error; D represents the experience pool.

[0193] Update the target network θ2 as follows:

[0194]

[0195] Among them, σ is the regularization parameter and ‖‖ is the regularization operation.

[0196] This embodiment simulates a simulation system for defending against advanced persistent threat (APT) attacks in an Industrial Internet of Things (IIoT) environment. The Industrial Internet of Things includes multiple legitimate IIoT device nodes and twenty honeypot nodes deployed in a communication network. The honeypot nodes are used to implement a defensive deceptive defense (DD) strategy by guiding the attacker to interact with the honeypot, thereby delaying or disrupting their attack path. After identifying the honeypot node, the attacker enters a new stage of attack and automatically exits the attack process after completing the key data collection. The responsible defender dynamically selects a response strategy for the honeypot node or the legitimate IIoT device node based on an in-depth understanding of the attacker's behavior pattern and a certain response delay.

[0197] The DD strategy is based on a multi-mechanism fusion design and incorporates hypergame modeling to simulate the interactions between attackers and defenders under conditions of information asymmetry. The simulation was built in Python 3.7 and implemented using PyTorch 1.10 as the deep learning framework. The experiment ran for 1000 rounds to evaluate the defensive effectiveness and system performance of different strategies in multi-round interactions.

[0198] Test indicators:

[0199] 1. Mean Time Between Failures (MTBF) measures the average operating time between failures, reflecting system stability and availability. MTBF helps evaluate the effectiveness of different DD strategies.

[0200] 2.FAR measures the average of false positives or detection rates, i.e. false alarm rate.

[0201] 3. BD-AHEU and BD-DHEU refer to the expected attack utility of the IIoT APT attacker based on the DD strategy and the expected defense utility of the defense middleware, respectively.

[0202] 4. To evaluate the performance of the proposed Collaborative Graph-guided Double Deep Q-Network (CG-D3QN) strategy, the system design includes a comparison of multiple adversarial strategy combinations. Participants can choose from three different strategy models: a strategy based on traditional game theory (TG), a strategy based on hypergame theory (HG), and a random strategy (RG). The defender can use the CG-D3QN strategy or a strategy based on a general machine learning (ML) algorithm in this model. The strategy execution process can further incorporate the decision of whether to enable or disable the defensive deception mechanism (Deceptive Defense (DD)).

[0203] By combining these variables, the system generates eight different strategy scenarios: HG-CG-D3QN-DD, TG-CG-D3QN-DD, HG-ML-DD, TG-ML-DD, HG-DD, TG-DD, RG-DD, and RG-ND. These strategy combinations are used to compare and evaluate the adaptability, robustness, and defensive effectiveness of CG-D3QN under various game participant strategies in a simulated environment, verifying its practical application in Industrial Internet of Things (IIoT) security defense.

[0204] The test results are as follows Figures 2 to 5 As shown, it is shown that the comprehensive evaluation of the super game defense system based on CG-D3QN provided in this embodiment has good adaptability, robustness and defense effect.

[0205] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An industrial Internet anti-deception technology system based on hyper-game, characterized in that: Includes industrial control layer, multi-access edge computing layer, and backend data layer; The industrial control layer is a distributed network including multiple nodes; the nodes are industrial Internet device nodes or honeypot nodes; the honeypot nodes are used to interfere with the attacker's identification of the real target and capture the characteristics of the attack behavior by creating a bait environment for virtual data interaction; The multi-access edge computing layer includes a plurality of edge computing nodes, wherein the edge computing nodes deploy anti-spoofing middleware for performing security authentication on data access requests; The backend data layer is used to store data required for the operation of the Industrial Internet and provide data services based on data access requests; The data access request raised by the node of the industrial control layer is securely authenticated by the node of the multi-access edge computing layer, and then the multi-access edge computing node provides data services to the node of the industrial control layer.

2. The industrial Internet anti-deception technology system based on hyper-game according to claim 1 is characterized in that: The edge computing node adopts a zero-trust mechanism for access to sensitive data. When the edge computing node receives a data access request, it formulates a current defense strategy and executes a defense action.

3. The industrial Internet anti-deception technology system based on hyper-game according to claim 1 is characterized in that: The defense middleware perceives the current attack stage of all attackers x in the hyper-game chain based on multi-source data and evaluates the uncertainty of the perception results. It adopts a decision model based on a partially observable Markov decision process to formulate the optimal defense strategy of the current hyper-game chain and execute defense actions.

4. The industrial Internet anti-deception technology system based on hyper-game according to claim 3 is characterized in that: The attack phases are based on MITRE's ATT&CK framework and include the following five: Initial access, which refers to the stage when an attacker first breaches the target network perimeter and establishes a foothold; Defense evasion refers to techniques used by attackers to conceal malicious behavior and bypass security detection; Credential access, which refers to obtaining user account passwords, tokens, or hashes to escalate privileges; Lateral movement, where an attacker spreads control within a target network and gains access to core data; Command control, where the attacker communicates with the controlled host through the server and issues instructions; Treat each attack phase as a subgame of the supergame.

5. The industrial Internet anti-deception technology system based on hyper-game according to claim 3 is characterized in that: The defense middleware perceives the current attack stage of the hyper game chain based on multi-source data, specifically: The defense middleware uses a machine learning classification algorithm or a rule engine classification to determine the current attack stage based on the detection results of the intrusion detection system, log analysis, and / or threat intelligence; wherein the intrusion detection system is used to detect known attack patterns or abnormal traffic; the log analysis specifically monitors system logs and network traffic data to identify abnormal behavior; the threat intelligence is attack behavior information updated by an external threat intelligence source.

6. The industrial Internet anti-deception technology system based on hyper-game according to claim 3 is characterized in that: The defense middleware evaluates the uncertainty of the perception result, specifically calculating the uncertainty of the perceived attacker x in the subgame h as follows Where ζ is a positive constant used to adjust the size and impact of the attacker's exposure to communication performance. Dx represents the degree of leakage of the defense deception strategy detected by the defense middleware. During the interaction process, the defense middleware simulates the attacker observing abnormal changes in system behavior and uses the ratio of the current environmental characteristic value to the expected environmental characteristic value without the defense deception strategy to represent the degree of leakage of the defense deception strategy. The environmental characteristic value is the weighted sum of the values of various environmental characteristics, including response time, data packet characteristics, and return information success rate. Attackers can exploit the leakage of the defense deception strategy and adjust their own attack strategy to deal with this defense mechanism, thereby introducing uncertainty in the defense middleware's perception. It represents the length of time the defense middleware observes the attacker. The longer the observation time, the smaller the perceived uncertainty. f(u) represents the impact of communication performance on the attacker. When the f(u) value is smaller, the communication stability is higher. It is defined as follows: f(u)=(1+BW n +DL+PL) / 3 Among them, BW n It is the bandwidth utilization rate of communication between the edge computing nodes of the multi-access edge computing layer and the industrial Internet device node n of the currently communicating industrial control layer; DL is the total transmission delay between the edge computing nodes of the multi-access edge computing layer and the industrial Internet device node n of the currently communicating industrial control layer; PL is the packet loss rate between the edge computing nodes of the multi-access edge computing layer and the industrial Internet device node n of the currently communicating industrial control layer.

7. The industrial Internet anti-deception technology system based on hyper-game according to claim 3 is characterized in that: The decision model based on partially observable Markov decision process is used to formulate the optimal defense strategy of the current hypergame chain, specifically: The defense middleware will calculate the attack stage h of all current attackers x and the perceived uncertainty As the current state, formulate the current defense strategy for node n based on the corresponding subgame h of the super game, and select the total expected defense utility value of the super game Minimal defensive strategy and execution of defensive moves; The total expected defense utility value of the super game is It can be evaluated by adding the attack utilities of all attackers in the Nash equilibrium state of subgame h.

8. The industrial Internet anti-deception technology system based on hyper-game according to claim 7 is characterized in that: The decision model based on partially observable Markov decision process is defined as follows: State S, including the communication performance indicators of the Industrial Internet, the perceived attack stage h, and the uncertainty of the attacker x in the defensive middleware perception subgame h The attack uncertainty of attacker x in a given subgame h Current state of defensive deception technology enabled d, system beliefs; Action space A, i.e., the set of defense strategies; Reward function r t , based on the expected utility, the Bellman equation is used to discount the future expected value of the reward and perform long-term optimal strategy planning, which is recorded as: Here, γ is a discount factor that is used to balance short-term and long-term rewards.

9. The industrial Internet anti-deception technology system based on hyper-game according to claim 8 is characterized in that: The communication performance indicators include the data transmission rate R of each link n (t), bandwidth utilization BW n , total transmission delay DL, packet loss rate Pl; System belief, which represents the true probability of adopting defense strategy t in the Industrial Internet neutron game h. Calculate as follows: in, The probability that the defensive intermediate adopts the defensive strategy t to respond to each node in the subgame h is, To defend the intermediate's initial belief about the environment and the opponent's behavior before the game begins, It refers to the total probability of responding by taking into account all possible defense strategies t; The attack uncertainty of attacker x in a given subgame h Calculate as follows: Among them, d is the defensive deception technology enabled status, which represents whether the system is using defensive deception technology to deal with potential attacks. When d is set to 1, it means that the system is actively using defensive deception technology to improve the speed and accuracy of response to attacks; when d is 0, it means that the system has not enabled this technology; It refers to the time that attackers monitor industrial control systems. The acquisition of monitoring time requires a combination of passive detection and active induction for perception.

10. The industrial Internet anti-deception technology system based on hyper-game according to claim 9 is characterized in that: Adopting a CG-D3QN network based on a D3QN network as the decision model based on the partially observable Markov decision process; The CG-D3QN network includes: state strategy network parameters The value function network θ1 is used to calculate the Q value and combine the loss function to back-propagate and train. The gradient descent method is often used to update the loss function. The value function network θ1 is used to predict the current Q value of the defense middleware. The target network θ2 is used to predict the Q value of the defense middleware at the next moment. The action advantage network Δ is used to judge the value difference of different actions under the current value function network θ1. And the state value network ψ is used to evaluate the current state S t the overall value of The action advantage network Δ and the state value network ψ are jointly updated using back propagation combined with gradient descent according to the Q-value error.

Citation Information

Patent Citations

  • ATT&CK-based spoofing defense system, construction method and full-link defense implementation method

    CN110430190A

  • Computing power network attack and defense confrontation game model based on supergame

    CN118821937A

  • Attack protection system based on RASP

    CN118827216A