Equipment collaborative security game defense method and system in industrial Internet of Things environment, terminal and storage medium
By constructing multi-dimensional device state vectors and collaborative feature matrices, and combining game theory and closed-loop self-evolution mechanisms, the defense strategy is dynamically adjusted, solving the problems of high false alarm rate and insufficient defense against complex attacks in the collaborative security defense of devices in the industrial Internet of Things environment, and achieving a balance between security and production at the millisecond level of real time.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 深圳开鸿数字产业发展有限公司
- Filing Date
- 2025-12-25
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, the security defense of device collaboration in the industrial IoT environment suffers from high false alarm rates and a lack of accurate modeling of the collaborative behavior of production line-level devices, resulting in insufficient defense against complex attacks. Furthermore, existing security solutions cannot effectively cope with complex collaborative attacks under millisecond-level real-time requirements.
By constructing a multi-dimensional device state vector and a collaborative feature matrix, the path probability distribution of system attackers is obtained. Combining game theory and a closed-loop self-evolution mechanism, the defense strategy is dynamically adjusted to achieve collaborative security defense for devices in the industrial Internet of Things environment.
It achieves a dynamic balance between security protection and production efficiency on a millisecond-level timescale, solving the problems of advanced threats being undefended, security measures being used incorrectly, and coordinated attacks being invisible, thus realizing the transformation from static passive protection to dynamic intrinsic security.
Smart Images

Figure CN121966947A_ABST
Abstract
Description
A method, system, terminal, and storage medium for device collaborative security game defense in an industrial Internet of Things (IoT) environment. Technical Field
[0001] This invention relates to the field of industrial Internet of Things (IoT) security, and in particular to a device collaborative security game defense method, system, terminal, and storage medium in an industrial IoT environment. Background Technology
[0002] Currently, with the rapid development of the Industrial Internet and intelligent manufacturing, industrial control systems are moving from closed, isolated systems to open, interconnected systems, facing unprecedented security challenges. Traditional boundary-based static protection (such as industrial firewalls and whitelists) is insufficient to cope with advanced persistent threats and complex collaborative attack chains within devices, and often conflicts with production efficiency and real-time requirements. Existing anomaly detection technologies have high false alarm rates and lack accurate modeling of collaborative behavior of production line-level equipment. General-purpose cybersecurity solutions cannot adapt to the diversity of industrial protocols and millisecond-level real-time control requirements, resulting in blind spots in security protection and creating a dilemma of "security versus production."
[0003] Therefore, there is an urgent need for a device collaboration security game defense method in the industrial Internet of Things environment to solve the problems of high false alarm rate of existing anomaly detection technology and insufficient defense against complex attacks due to lack of accurate modeling of the collaborative behavior of production line-level devices.
[0004] Therefore, existing technologies still need to be improved and enhanced. Summary of the Invention
[0005] To address the aforementioned shortcomings of existing technologies, this invention provides a method, system, and terminal for defending against collaborative security games in an industrial Internet of Things (IoT) environment. The aim is to solve the problems of high false alarm rates in existing anomaly detection technologies and insufficient defense against complex attacks due to a lack of accurate modeling of collaborative behavior among production line-level devices.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a device collaborative security game-theoretic defense method in an industrial Internet of Things (IoT) environment. The method includes: acquiring target data; constructing a target multi-dimensional device state vector and a target collaborative feature matrix based on the target data, wherein the target data includes the operating data of each device on the target production line; constructing a role-action space mapping; constructing a target payoff matrix based on the role-action space mapping; and obtaining the target path probability distribution of a system attacker based on the target payoff matrix; obtaining target anomaly deviation and target synchronization decay factor based on the target multi-dimensional device state vector and the target collaborative feature matrix; constructing a target security offset index formula based on the target anomaly deviation and target synchronization decay factor, wherein the target anomaly deviation and target synchronization decay factor are respectively the anomaly deviation and decay factor of the devices on the target production line; obtaining a target defense strategy based on the target security offset index formula and the target path probability distribution; and dynamically defending the target production line based on the target defense strategy.
[0007] In one implementation, the target multidimensional device state vector is: ;in, Let be the state vector of device d at time t. Represents matrix transpose. This represents the PLC instruction sequence for device d. Indicates from A function to extract instruction execution frequency characteristics; Let d be the set of communication messages of device d. for A function for calculating information entropy; Let d be the response time series of device d. From A function to extract the delay fluctuation characteristics.
[0008] In one implementation, the target collaborative feature matrix is: ;in, Indicates equipment and equipment Dependency relationship between them express and Statistical dependence between them For equipment State vector, For equipment State vector; This indicates the synchronization relationship of all equipment on the target production line. This represents the set of all equipment on the target production line. for The momentum, for The orientation angle in three-dimensional space, for The average value of the orientation angle of the state vectors of all devices in the system.
[0009] In one implementation, the role-action space mapping is as follows: ;in, This provides the action space for system attackers. Let be the strategy probability vector of the attacker in the system; For the action space of the system defenders, Let be the system defender strategy probability vector.
[0010] In one implementation, the target return matrix is: ; ; ;in, For the utility of the system defender, The weighting coefficients and , For the safety scoring function, This represents the current system state. The actions of the attacker in the system. For the actions of the system defenders, for Production losses; For the safety scoring formula, The current state vector With baseline state vector The Euclidean distance between them This is the maximum acceptable deviation distance; The formula for production loss rate is as follows: To execute The delay time, This is the initial cycle time.
[0011] In one implementation, the target path probability distribution is: ; System status Under these conditions, the attacker selects an action in the system. The conditional probability; For a temperature parameter greater than 0, Indicates the attacker is in state Next, take action And the system defenders take the optimal response. At that time, the gains obtained by the attacker of the system, This represents any attack action by the attacker in the system; argmax is the parameter used to make the following function reach its maximum value.
[0012] In one implementation, the target safety offset index formula is: ; ; ;in, Indicates device At any moment The aforementioned target anomaly deviation, Indicates device The current state vector, Indicates device The mean vector of the historical normal state, Indicates device Covariance matrix of historical normal state vector The inverse matrix; This represents the target synchronization attenuation factor of the target production line. For the target production line at time The actual synchronization rate This refers to the baseline synchronization rate of the target production line under historical normal operating conditions. Indicates all individual devices Find the average value. Indicates the target production line Find the average value.
[0013] In one implementation, obtaining a target defense strategy based on the target security offset index formula and the target path probability distribution includes: obtaining a preset defense threshold, the preset defense threshold including a first threshold and a second threshold, the first threshold being used for conventional production lines, the second threshold being used for high-precision manufacturing production lines, and the second threshold being less than the first threshold; and obtaining the target defense strategy based on the preset defense threshold.
[0014] In one implementation, obtaining the target defense strategy based on the preset defense threshold includes: constructing a hierarchical strategy decision tree based on the target security offset index formula and the preset defense threshold; obtaining the target risk level based on the hierarchical strategy decision tree, wherein the target risk level includes low risk, medium risk, and high risk; and obtaining the target defense strategy based on the target risk level.
[0015] In one implementation, the hierarchical strategy decision tree is: ;in, The preset defense threshold, Represents the policy decision function. The aforementioned low risk; The aforementioned medium risk; This is considered a high-risk situation.
[0016] In one implementation, obtaining the target defense strategy based on the target risk level includes: if the target risk level is low risk, the target defense strategy is to only log; if the target risk level is medium risk, the target defense strategy is encryption and authentication; if the target risk level is high risk, the target defense strategy is game theory active defense.
[0017] In one implementation, before obtaining the target defense strategy, the method further includes: detecting the current capacity utilization rate of the target production line; if the current capacity utilization rate is higher than the target threshold, then increasing the preset defense threshold by 0.05.
[0018] In one implementation, obtaining the target defense strategy based on the target risk level includes: when the target risk level is high risk, obtaining a game-theoretic optimization formula, and optimizing the target defense strategy based on the game-theoretic optimization formula; the game-theoretic optimization formula is: ;in, This indicates that the predicted attack actions against the attacker targeting the system are taken at the maximum value. This indicates that the defensive action taken by the defender against the system is the minimum value. This indicates the gains of the attacker in the system. Benefits of the system defender The expected value of the difference For the actions of the system defender The impact on the production of the target production line is greater than or equal to 0.
[0019] In one implementation, obtaining the target defense strategy based on the target risk level further includes: constructing a strategy deployment function based on the target path probability distribution, and implementing the game-theoretic active defense strategy based on the strategy deployment function; the strategy deployment function is: ;in, For activation function, For the secure kernel of the OpenHarmony operating system, This represents the probability distribution vector of the target defense strategy. Represents a timestamp.
[0020] In one implementation, before dynamically defending the target production line based on the target defense strategy, the method further includes: constructing a closed-loop feedback mechanism to dynamically update the preset defense threshold through gradient optimization based on the actual security effect and production impact of historical defense strategies.
[0021] In one implementation, dynamically updating the preset defense threshold using gradient optimization includes: dynamically updating the preset defense threshold based on a target update formula; the target update formula is: ;in, Indicates the first The preset defense threshold for the next iteration Indicates the first The preset defense threshold for the next iteration Indicates the learning rate. This represents the overall objective function that maximizes profits. Indicates to Find the partial derivatives.
[0022] A second aspect of the present invention provides a device collaborative security game defense system in an industrial Internet of Things (IoT) environment, comprising: a construction module for acquiring target data, constructing a target multi-dimensional device state vector and a target collaborative feature matrix based on the target data, wherein the target data includes the operating data of each device on the target production line; a collaborative modeling module for constructing a role-action space mapping, constructing a target payoff matrix based on the role-action space mapping, and obtaining a target path probability distribution of the system attacker based on the target payoff matrix; a security quantification module for obtaining a target abnormal deviation and a target synchronization decay factor based on the target multi-dimensional device state vector and the target collaborative feature matrix, constructing a target security offset index formula based on the target abnormal deviation and the target synchronization decay factor, wherein the target abnormal deviation and the target synchronization decay factor are respectively the abnormal deviation and decay factor of the devices on the target production line; and a strategy acquisition module for obtaining a target defense strategy based on the target security offset index formula and the target path probability distribution, and performing dynamic defense on the target production line based on the target defense strategy.
[0023] A third aspect of the present invention provides a terminal, the terminal including a processor and a computer-readable storage medium communicatively connected to the processor, the computer-readable storage medium being adapted to store a plurality of instructions, the processor being adapted to invoke the instructions in the computer-readable storage medium to execute the steps of implementing the device collaborative security game defense method in the industrial Internet of Things environment described in any of the preceding claims.
[0024] In a fourth aspect, the present invention provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the device collaborative security game defense method in the industrial Internet of Things environment described in any of the preceding claims.
[0025] Compared with existing technologies, this invention provides a device collaborative security game-theoretic defense method, system, and terminal in an industrial Internet of Things (IIoT) environment. The method involves acquiring target data, constructing a target multi-dimensional device state vector and a target collaborative feature matrix based on the target data (including operational data of each device on the target production line), then constructing a role-action space mapping, building a target payoff matrix based on the mapping, and obtaining the target path probability distribution of the system attacker based on the payoff matrix. Next, based on the target multi-dimensional device state vector and the target collaborative feature matrix, target anomaly deviation and target synchronization decay factor are obtained. A target security offset index formula is constructed based on the target anomaly deviation and target synchronization decay factor, where the target anomaly deviation and target synchronization decay factor are respectively the anomaly deviation and decay factor of the devices on the target production line. Finally, a target defense strategy is obtained based on the target security offset index formula and the target path probability distribution, and dynamic defense is performed on the target production line based on the target defense strategy. The device collaborative security game-theoretic defense method proposed in this invention for the industrial Internet of Things (IIoT) environment solves the problems of high false alarm rates and insufficient defense against complex attacks caused by the lack of accurate modeling of collaborative behavior of production line-level equipment in existing technologies. By constructing an intelligent defense system that integrates real-time behavioral profiling, game theory-based proactive decision-making, and closed-loop self-evolution capabilities, it dynamically balances security protection and production efficiency on a millisecond-level timescale. This fundamentally solves the core contradictions of advanced threats being undefended, security measures leading to erroneous production, and collaborative attacks being invisible in industrial environments, achieving a paradigm shift from static passive protection to dynamic intrinsic security. Attached Figure Description
[0026] Figure 1 is a flowchart of an embodiment of the device collaborative security game defense method in an industrial IoT environment provided by the present invention; Figure 2 is an overall technical flowchart of an embodiment of the device collaborative security game defense method in an industrial IoT environment provided by the present invention; Figure 3 is a flowchart of the first step in obtaining the defense strategy of an embodiment of the device collaborative security game defense method in an industrial IoT environment provided by the present invention; Figure 4 is a flowchart of the second step in obtaining the defense strategy of an embodiment of the device collaborative security game defense method in an industrial IoT environment provided by the present invention; Figure 5 is a flowchart of the system decision-making process of an embodiment of the device collaborative security game defense method in an industrial IoT environment provided by the present invention; Figure 6 is a flowchart of the industrial scenario verification process of an embodiment of the device collaborative security game defense method in an industrial IoT environment provided by the present invention; Figure 7 is a structural principle diagram of an embodiment of the device collaborative security game defense system in an industrial IoT environment provided by the present invention; Figure 8 is a schematic diagram of the principle of an embodiment of the terminal provided by the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0028] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0029] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0030] The device collaborative security game defense method in the industrial Internet of Things (IoT) environment provided by this invention can be applied to terminals with computing capabilities. The terminals can execute the device collaborative security game defense method in the industrial IoT environment provided by this invention to ensure the security of industrial IoT devices.
[0031] This embodiment describes a device-based collaborative security game-theoretic defense method in an industrial Internet of Things (IIoT) environment. In this embodiment, by constructing an industrial intrinsic security intelligent agent based on digital profiling, game theory, and adaptive decision-making, precise defense against advanced threats and an autonomous balance of security effectiveness are achieved.
[0032] Specifically, in the current wave of deep integration between the Industrial Internet and intelligent manufacturing, the security protection of industrial control systems is facing unprecedented and severe challenges. Existing technologies primarily rely on a combination of "boundary protection" and "static rules," manifested as follows: deploying industrial firewalls at network boundaries to perform deep packet inspection of industrial control protocols (such as Modbus TCP and Profinet); employing application whitelisting mechanisms at the controller level to allow only pre-authorized commands to execute; and a central security information and event management platform for log aggregation and alerting. The current technical characteristics of this system are reliance on prior knowledge, fixed rules, and delayed response. Its core operational logic lies in matching known attack characteristics with traffic patterns, and once a match is found, blocking or alerting is implemented.
[0033] However, existing technologies face three critical and urgent technical challenges. First, static defense mechanisms are completely ineffective against advanced attacks. Modern attacks against industrial systems increasingly employ legitimate command formats for malicious operations. For example, sending compliant but high-frequency commands to a welding robot to cause it to accelerate (to 150% of its rated power) is completely ineffective against traditional firewalls and whitelists, as they cannot understand the semantics and context of the commands. Statistics show that such command-level attacks result in a false positive rate exceeding 35%, because fluctuations in equipment during normal operating conditions such as startup and model changeovers are incorrectly identified as attacks by static thresholds, severely disrupting production. Second, there is an irreconcilable contradiction between security measures and production efficiency. The core requirements of industrial production are determinism and continuity, but existing security solutions severely violate these principles. For example, end-to-end encryption implemented to ensure communication security extends the production line control cycle by 22% to 30%; continuous cross-authentication between devices consumes an additional 18% of the terminal's total power consumption. This leads to production departments often disabling or weakening security functions for efficiency reasons in actual deployments, creating significant security vulnerabilities. Third, there are blind spots in the protection of cross-device collaborative attack chains. Modern intelligent production lines rely on the close collaboration of devices such as PLCs, robotic arms, and sensors. Attackers can exploit the trust relationships between these devices to construct complex attack paths. For example, they could compromise a PLC and then use it as a springboard to send malicious commands to the robotic arm. Existing isolated security devices, such as standalone PLC protection modules, lack the ability to perceive and coordinate global collaborative behavior, making them completely "invisible" to such attack chains. Furthermore, the heterogeneity of communication protocols between devices from different manufacturers makes it difficult to deploy a unified protection strategy synchronously.
[0034] To address the aforementioned issues, existing technologies employ machine learning-based anomaly detection models to enhance detection capabilities, analyzing network traffic or device logs to identify deviations from the baseline. However, these improvements suffer from drawbacks: the models are typically "black boxes," lacking interpretability and struggling to adapt to rapidly changing industrial environments, resulting in persistently high false alarm rates. To resolve the conflict between security and production, some existing technologies propose manual switching between "security mode" and "production mode," but this relies on manual operation, leading to slow response times and a high risk of errors, ultimately failing to achieve dynamic balance. To counter coordinated attacks, some research attempts to centrally distribute policies through a unified management platform; however, limited by network latency and protocol conversion overhead, policy synchronization delays often exceed 150 milliseconds, failing to meet the stringent millisecond-level real-time synchronization requirements of high-end manufacturing (such as semiconductor lithography equipment clusters requiring phase deviation tolerance of less than 5°). These improvements either introduce new complexities or fail to fundamentally solve the problems, exposing the inherent limitations of existing technological architectures in addressing intelligent and collaborative advanced threats.
[0035] Based on this, this embodiment proposes a device collaborative security game-theoretic defense method in an industrial IoT environment. The proposed method has wide and critical applications in various industrial IoT environments. Specifically, in the field of high-end manufacturing production line safety control, it can achieve real-time interception of abnormal commands from automotive manufacturing welding robotic arm clusters within 10 milliseconds, ensuring precise calibration of the collaborative work cycle of semiconductor lithography equipment clusters; in critical infrastructure IoT protection, it can dynamically deploy encrypted links between substation smart terminals and verify the integrity of urban gas pipeline network sensor data; in cross-device collaborative operation environments, it can provide safety command arbitration for port unmanned cranes and AGVs, and dynamically allocate operating permissions for medical surgical robot clusters. Its typical use cases are highly targeted: when the PLC controller receives more than 100 high-frequency commands within 0.1 seconds, the system can immediately trigger a game-theoretic defense sandbox to perform attack simulations; it can automatically update the behavioral baseline template during the maintenance period of the device cluster to adapt to parameter drift caused by equipment wear; when a new node joins the network, a distributed arbitration module can achieve seamless and rapid synchronization of security policies.
[0036] Specifically, as shown in Figure 1, in one embodiment of the device collaborative security game defense method in the industrial Internet of Things environment provided by the present invention, the device collaborative security game defense method in the industrial Internet of Things environment includes the following steps: S100, acquiring target data, constructing a target multi-dimensional device state vector and a target collaborative feature matrix based on the target data, wherein the target data includes the operating data of each device on the target production line.
[0037] Specifically, referring to Figure 2, in this embodiment, firstly, the target multidimensional equipment state vector and the target collaborative feature matrix are constructed based on the dynamic baseline of the behavior of each device on the target production line. Specifically, the target data can be extracted from the behavior monitoring of the corresponding industrial equipment.
[0038] Then, by using three-dimensional modeling of instruction frequency, communication entropy, and response delay, a digital fingerprint of the device is established, namely the target multi-dimensional device state vector. At the same time, the target collaborative feature matrix is constructed, laying the foundation for solving the problem that static rule bases are difficult to adapt to production line operating condition fluctuations.
[0039] Specifically, the equipment state vector modeling formula on the target production line, i.e., the target multidimensional equipment state vector, is as follows: ;in, For equipment The state vector at time t, where, It is a vector, that is, a quantity with direction, and its subscript is... Indicates the equipment number, superscript Indicates time, Represents matrix transpose. This represents the PLC instruction sequence for device d. Indicates from A function to extract instruction execution frequency characteristics; This is the set of communication packets for device d. for The function that calculates information entropy is used to measure the uncertainty or randomness of information. This is the response time series of device d. From A function to extract the delay fluctuation characteristics.
[0040] In this embodiment, instruction frequency characteristics It is used to analyze the instruction sequence of a PLC within a very short time window, which can be: Information entropy formula used to detect abnormal high-frequency operations (such as frequent start-stop); communication entropy value characteristics. It is the information entropy of communication messages between computing devices. Normal communication has a certain degree of randomness, meaning its entropy value is relatively high. However, encryption interference or fixed attack payloads can cause the communication pattern to solidify, resulting in a lower entropy value. In this embodiment, a sudden drop in entropy value >30% is considered a risk. The formula for information entropy is: ;in, Represents the set of communication messages Information entropy Indicate message type The probability of occurrence in a communication stream. This represents a logarithm to the base 2. It is commonly used in information theory, and its unit is bit.
[0041] The information entropy formula is used to calculate an index of the randomness of communication messages. The more uniform (more chaotic) the probability distribution, the higher the entropy value; the more concentrated (more regular) the probability distribution, the lower the entropy value. Specifically, in this embodiment, the communication entropy characteristics... Information entropy used to calculate device communication messages Normal communication has a certain degree of randomness (high entropy value), while encryption interference or fixed attack payloads can cause communication patterns to solidify, and a sudden drop in entropy value greater than 30% is considered a risk.
[0042] Response latency characteristics The standard deviation of the response delay is usually taken. Dramatic fluctuations in latency may indicate that the system has been slowed down or is experiencing interference.
[0043] The target collaborative feature matrix is: ;in, Indicates equipment and equipment , equipment and equipment Dependency relationship between them express and The statistical dependency between them specifically refers to the statistical dependency between the states of two devices quantified using mutual information. A larger value indicates a stronger behavioral correlation between the two devices. Furthermore... For equipment State vector, For equipment State vector; This indicates the synchronization relationship of all equipment on the target production line; sync means synchronization. This represents the set of all equipment on the target production line. for The potential, i.e., the set The number of elements in the middle. for The orientation angle in three-dimensional space, for The average value of the orientation angle of the state vectors of all devices in the system. This is used to calculate the cosine value of each device in the cluster relative to the overall average direction, and then take the average. The closer the cosine value is to 1, the more consistent the directions among the devices, and the better the synchronization.
[0044] Specifically, dependency measurement Used for mutual information This measures the statistical correlation between the state vectors of two devices. An abnormally high correlation (e.g., being controlled by the same attacker) or a abnormally low correlation (e.g., a failure of coordination) indicates an anomaly.
[0045] Synchronization state measurement Collection of computing devices Angle of all device state vectors with average angle The cosine mean of the phase difference. An alarm is triggered if the phase difference is greater than 15 degrees, indicating that the production line cycle is out of order.
[0046] In step S100, through real-time calculation and This is equivalent to the system building a dynamic "digital profile" and "collaborative baseline" for each device and each production line, serving as a benchmark for judging whether everything is normal.
[0047] Referring again to Figure 1, in this embodiment, the device collaborative security game defense method in the industrial Internet of Things environment further includes the following steps: S200, constructing a role-action space mapping, building a target reward matrix based on the role-action space mapping, and obtaining the target path probability distribution of the system attacker based on the target reward matrix.
[0048] The role-action space mapping is as follows: ;in, This provides the action space for system attackers. Let be the strategy probability vector of the attacker in the system; For the action space of the system defenders, Let be the system's defender strategy probability vector. This is a set of symbols used to list three specific actions that a system attacker or defender can take. In more embodiments, it may include more specific actions.
[0049] In this mapping relationship, the attacker is explicitly defined. and defenders The set of optional actions is mapped to a specific policy probability vector. .
[0050] The target return matrix is as follows: ; ; ;in, For the utility or benefit of the system defender, is the weighting coefficient, which is a real number greater than 0, where , For the safety scoring function, This represents the current system state. The actions of the attacker in the system. For the actions of the system defenders, for Production losses; For the safety scoring formula, The current state vector With baseline state vector The Euclidean norm between them, i.e., the Euclidean distance. (Subscript) Represents the L2 norm. This is the preset maximum acceptable deviation distance. Overall, the greater the distance, the lower the safety score. The standardized distance is subtracted from 1, resulting in a score between 0 and 1. The formula for production loss rate is as follows: To execute The additional delay time introduced. This is the initial cycle time, which is the standard cycle time of the production process.
[0051] Specifically, the effectiveness of the defender It is the weighted difference between safety gains and production losses. This reflects the core contradiction in industrial scenarios—the balance between safety and efficiency.
[0052] A security score was constructed based on the target benefit matrix. Based on the state vector from the first part. Current state. Deviation from baseline The smaller the Euclidean distance, the higher the safety score, with a maximum safety score of 1; furthermore, a production loss rate was also constructed. Specifically refers to defensive actions Processing delays caused by (such as encryption and isolation) Production cycle The proportion; furthermore, weights were also constructed. This reflects the decision-maker's safety preferences. As can be seen in this embodiment, the target benefit matrix reflects that the current production line places greater emphasis on safety.
[0053] The probability distribution of the target path is as follows: ; System status Under these conditions, the attacker selects an action in the system. The conditional probability; It is a temperature parameter greater than 0, used to control the sharpness of the probability distribution. The larger the value, the higher the probability that a high-yield action will be selected. Indicates the attacker is in state Next, take action And the system defenders take the optimal response. At that time, the gains obtained by the attacker of the system, This represents any attack action by an attacker against the system. The summation symbol in the denominator represents the summation over all possible attack actions. Summing ensures that the sum of all probabilities is 1, where e is the natural constant, approximately equal to 2.718; For the optimal response strategy of the system's defenders, argmax is the parameter that maximizes the value of the subsequent function.
[0054] The target path probability distribution is a Softmax function that transforms the attacker's payoff (utility) into a probability distribution, followed by constraints. This refers to a specific action taken by the attacker. The defender's optimal response It's the one that maximizes its own defensive benefits. The action.
[0055] In other words, in this embodiment, it is assumed that the system attacker is "rational" and chooses different attack actions. The probability follows a Softmax distribution, which is determined by the attack payoff. Decide.
[0056] The key to this step is that, when making predictions, the attacker assumes that the defender will take the optimal defensive action against the current attack. This is a nested optimization process that ultimately solves for the Nash equilibrium—a stable state where neither side can benefit from unilaterally changing its strategy.
[0057] Specifically, in this embodiment, an offensive and defensive digital twin is constructed to simulate adversarial combat in a virtual environment. By solving the game equilibrium, the system can predict high-probability attack paths (>85% probability hotspots) and provide the defender with the optimal response strategy under the equilibrium state. .
[0058] S300. Based on the target multidimensional equipment state vector and the target collaborative feature matrix, obtain the target abnormal deviation and the target synchronization attenuation factor. Based on the target abnormal deviation and the target synchronization attenuation factor, construct the target safety offset index formula, where the target abnormal deviation and the target synchronization attenuation factor are the abnormal deviation and attenuation factor of the equipment on the target production line, respectively.
[0059] In this step, complex multi-dimensional monitoring data, the target multi-dimensional device state vector, and the target collaborative feature matrix are aggregated into a simple security posture index, which is used to trigger different levels of defense.
[0060] The formula for the target safety offset index is: ; ; ;in, Indicates equipment At any moment The aforementioned target anomaly deviation, Indicates equipment The current state vector, Indicates equipment The mean vector of the historical normal state, Indicates equipment Covariance matrix of historical normal state vector The inverse matrix; This represents the target synchronization attenuation factor of the target production line. For the target production line at time The actual synchronization rate, This is the baseline synchronization rate of the target production line under historical normal operating conditions; Indicates all individual devices Find the average value. Indicates the target production line Find the average value.
[0061] in, Mahalanobis distance is used to measure the status of a single device. Deviating from its historical normal pattern (mean is The covariance matrix is The multivariate comprehensive distance is calculated by considering the correlation between features, making it more accurate than the Euclidean distance. In this embodiment, A value greater than 2.5σ will directly classify the corresponding device as a high-risk node.
[0062] The target synchronous decay factor of the target production line This is used to quantify the entire group of equipment in the target production line. Current synchronization rate Synchronization rate with historical normal The absolute difference. In this embodiment, a level three warning will be triggered directly when the synchronization rate decay is greater than 10%.
[0063] Furthermore, It is a comprehensive safety risk score for the entire system. It is a weighted average of the anomaly level of all equipment with a weight of 0.6, and a weighted average of the synchronous attenuation level of all production lines with a weight of 0.4.
[0064] Furthermore, in this embodiment, a preset defense threshold is also constructed, which is also the threshold for triggering defense. . Specifically, The safety tolerance is dynamically adjusted based on the production scenario. The threshold for chip manufacturing (high precision) is more stringent than that for conventional assembly (conventional production lines) (0.1 vs 0.2). In other words, for conventional production lines, the preset defense threshold... Set to 0.2; for high-precision manufacturing (such as chips), the preset defense threshold is... Then a stricter threshold of 0.1 is adopted.
[0065] S400. Obtain a target defense strategy based on the target security offset index formula and the target path probability distribution, and perform dynamic defense on the target production line based on the target defense strategy.
[0066] Specifically, the target safety offset index formula is calculated as follows: The value reflects the real-time, quantitative, and comprehensive safety risk level currently faced by the production line. Simultaneously, the target path probability distribution is output by the game theory deduction module. It accurately depicts the various attack actions that an attacker might take under the current system state. The probability of (such as falsified data, overclocking, etc.). This step focuses on reflecting the level of risk. The system intelligently integrates and maps the probability distribution of attack paths ("how the attack will occur") to generate an optimized set of defense instructions that most effectively responds to the most likely attacks in a specific risk scenario—the target defense strategy. Finally, this strategy is deployed in milliseconds to relevant PLC controllers, robotic arms, sensors, and other target devices on the production line via the industrial operating system kernel. This achieves real-time, precise, and adaptive dynamic security defense for the production line, ensuring maximum continuity and stability of production while mitigating threats. Specifically, in this embodiment, the target defense strategy is deployed via a distributed bus based on the OpenHarmony system.
[0067] Referring to Figure 3, the target defense strategy is obtained based on the target security offset index formula and the target path probability distribution, including: S410, obtaining a preset defense threshold, wherein the preset defense threshold includes a first threshold and a second threshold, the first threshold is used for conventional production lines, the second threshold is used for high-precision manufacturing production lines, and the second threshold is less than the first threshold.
[0068] Specifically, in this embodiment, a configurable defense trigger threshold closely related to the security tolerance of the production scenario is preset. The preset defense threshold includes at least two key values: namely, the first threshold... and the second threshold Among them, the first threshold For conventional production lines with relatively high tolerance for production interruptions and relatively relaxed requirements for process precision (e.g., general component assembly lines, conventional packaging lines), a typical empirical value can be set as follows: =0.2. The second threshold This is applied to high-precision manufacturing lines with extremely stringent requirements for production continuity, stability, and process accuracy, such as chip lithography production lines, aero-engine precision machining lines, and high-purity biopharmaceutical production lines. Typical empirical values for these lines are set to be even more stringent. =0.1, that is < This differentiated threshold setting stems from the inherent safety requirements of different industrial scenarios: in high-precision manufacturing, even minute anomalies, such as a 5° phase deviation or millisecond-level timing jitter, can lead to the scrapping of an entire batch of products. Therefore, a more sensitive threshold must be adopted, that is... Lower threshold values enable earlier and more stringent defense responses; while for conventional production lines, thresholds can be appropriately relaxed while ensuring basic safety to reduce false alarms and unnecessary defense actions caused by normal production fluctuations, thus achieving a more suitable balance between safety and efficiency. The specific threshold values can be automatically matched according to the production line type during system initialization, or fine-tuned by the safety administrator based on historical data and risk assessment.
[0069] S420. Obtain the target defense strategy based on the preset defense threshold.
[0070] Specifically, in this embodiment, the calculated target safety offset index will be... , and the preset defense threshold determined according to the current production line type ( or The system performs real-time comparisons and, based on a set of pre-defined, tiered, and progressive logical rules, maps and triggers a corresponding set of specific defensive action commands. This process is not a simple "yes / no" judgment, but a "strategy upgrade" process based on risk level and attack prediction, moving from passive monitoring to proactive countermeasures. Its purpose is to ensure that the defense strength is precisely matched with the risk level, avoiding insufficient or excessive defense.
[0071] Referring to Figure 4, the target defense strategy is obtained based on the preset defense threshold, including: S421, constructing a hierarchical strategy decision tree based on the target security offset index formula and the preset defense threshold, and obtaining the target risk level based on the hierarchical strategy decision tree, wherein the target risk level includes low risk, medium risk and high risk.
[0072] The hierarchical strategy decision tree is as follows: ;in, The preset defense threshold, Represents the policy decision function. The aforementioned low risk; The aforementioned medium risk; This is considered a high-risk situation.
[0073] Specifically, referring to Figure 5, the hierarchical strategy decision tree is a predefined, condition-triggered automated decision logic function, denoted as... ,in For the applicable preset defense threshold It can be or This decision tree will use consecutive... The numerical value is compared with a discrete threshold range to divide the current global security situation into three distinct target risk levels: Low risk: if and only if This level indicates that the overall operating status of the production line is highly consistent with the historical normal baseline, no significant coordination anomalies have been detected, and the system believes that the probability of facing a substantial attack threat is extremely low.
[0074] The aforementioned medium risk: when Time-based determination. This level indicates that a certain degree of individual device anomaly or cluster synchronization decay has been detected, the security posture has clearly shifted, there are potential vulnerabilities that can be exploited, and the probability of an attack has significantly increased.
[0075] The high risk mentioned: when This level indicates that a severe or multi-point concurrent anomaly has been detected, suggesting that production line collaboration may have been compromised. The system determines that the line is facing or about to face a high-probability, high-impact, substantial cyberattack, placing both production and physical security under imminent threat. Through this decision tree, the system achieves a high degree of abstraction and simplification of complex, multi-dimensional security data, outputting a clear and actionable "risk signal" that provides a fundamental basis for subsequently selecting specific defense measures.
[0076] S422. Obtain the target defense strategy based on the target risk level.
[0077] The target defense strategy is obtained based on the target risk level, including: if the target risk level is low risk, the target defense strategy is to only log; if the target risk level is medium risk, the target defense strategy is encryption and authentication; if the target risk level is high risk, the target defense strategy is game theory active defense.
[0078] Specifically, after determining the target risk level, predefined defense strategy modules are activated, and the target path probability distribution is considered. The strategy is fine-tuned to ultimately generate the target defense strategy.
[0079] In this embodiment, the strategy paradigms corresponding to each risk level are as follows: If the target risk level is low risk, the target defense strategy is to only log, and in this case, the monitoring and recording mode is activated. The target defense strategy mainly involves continuously recording all security-related logs and performance data, updating the behavioral baseline model, but without implementing any proactive blocking or encryption actions that may interfere with the production process. This mode has almost zero impact on production and aims to maintain situational awareness at the lowest cost.
[0080] If the target risk level is classified as medium risk, the target defense strategy is encryption and authentication. In this case, the basic enhanced defense mode is activated. The target defense strategy will include a series of lightweight but effective enhancements, such as: dynamically enabling link encryption (rather than full-link) for critical control command flows, triggering one-time cross-authentication between critical devices, and rate limiting the communication rate of devices suspected of being anomalous sources. These strategies aim to increase the cost and difficulty for attackers, curb risk escalation, and their design goal is to keep the average impact on the production cycle at a low level.
[0081] If the target risk level is high risk, the target defense strategy is game-theoretic active defense. In this case, the game-theoretic active defense mode is activated. This is the core advanced response of this embodiment. The generation of the target defense strategy will no longer rely solely on fixed rules, but will instead incorporate the current state. The target path probability distribution and the preset attack and defense benefit matrix The input is fed into the game theory optimization engine to solve for the Nash equilibrium or optimal response strategy under the current constraints. The generated strategies may include: immediately isolating nodes identified as high-probability attack sources or compromised nodes; dynamically reconstructing parts of the network topology or security domains; and executing pre-defined emergency security scripts. This mode aims to maximize security benefits, although it may have a high temporary production impact, but this is a necessary action to prevent catastrophic production accidents or equipment damage. In this mode, the target path probability distribution is directly used to locate the highest-risk attack paths, enabling defense resources to be precisely deployed to the most critical links.
[0082] Ultimately, through hierarchical decision-making, an automated closed loop is achieved, from risk perception to level determination to strategy matching to precise execution, ensuring dynamic, efficient, and intelligent safety protection for industrial production lines.
[0083] Further, obtaining the target defense strategy based on the target risk level includes: when the target risk level is high risk, obtaining a game-theoretic optimization formula, and optimizing the target defense strategy based on the game-theoretic optimization formula; the game-theoretic optimization formula is: ;in, This indicates that the predicted attack actions against the attacker targeting the system are taken at the maximum value. This indicates that the defensive action taken by the defender against the system is the minimum value. This indicates the gains of the attacker in the system. Benefits of the system defender The expected value of the difference For the actions of the system defender The impact on the production of the target production line is greater than or equal to 0.
[0084] Specifically, in high-risk mode, the goal of the system defender shifts to: satisfying production constraints. Under the premise of (e.g., minimum throughput), minimize the gap between the attacker's maximum expected gain and that of oneself.
[0085] Furthermore, obtaining the target defense strategy based on the target risk level further includes: constructing a strategy deployment function based on the target path probability distribution, and implementing the game-theoretic active defense strategy based on the strategy deployment function; the strategy deployment function is: ;in, For activation function, This is the security kernel of the OpenHarmony operating system, and is the software module that specifically executes the defense strategy in this embodiment. This represents the probability distribution vector of the target defense strategy. Represents a timestamp.
[0086] Specifically, at this point, the calculated optimal defense strategy will be... (or strategy distribution) The control is executed on specific industrial equipment at a specified timestamp via the secure kernel of the OpenHarmony(OH) operating system. In this embodiment, a delay of less than 2ms is required to ensure real-time control.
[0087] Thus, in this embodiment, a revolutionary breakthrough in industrial security protection, from decision-making to execution, is achieved through multi-level collaborative optimization. The entire system is built upon three pillars: dynamic trade-offs, hardware acceleration, and deep integration with the operating system.
[0088] In this embodiment, before obtaining the target defense strategy, the method further includes: detecting the current capacity utilization rate of the target production line; if the current capacity utilization rate is higher than the target threshold, then increasing the preset defense threshold by 0.05.
[0089] Specifically, before acquiring the target defense strategy, a capacity-aware adaptive adjustment step is also included.
[0090] Specifically, before making a final strategy decision through the hierarchical strategy decision tree, the system will monitor the current capacity utilization rate of the target production line in real time. The current capacity utilization rate is calculated by collecting data from the production line's overall control system, such as a Manufacturing Execution System (MES) or Supervisory Control and Data Acquisition (SCADA), and comparing the actual output quantity with the theoretical maximum design capacity within the most recent statistical period (e.g., 15 minutes).
[0091] In this embodiment, a target threshold for capacity utilization is preset, which can be 95% in this embodiment. This threshold represents a critical production load threshold. Exceeding this point means that the production line is operating at full or overload under high pressure. Any additional delays or interruptions may lead to major production accidents such as order delivery defaults, work-in-process backlogs, or production rhythm disruptions.
[0092] Specifically, it can be seen that the device collaborative security game defense method in the industrial IoT environment proposed in this embodiment has the ability to automatically degrade strategies. Its principle is dynamic risk-capacity equilibrium adjustment. When the production line capacity utilization rate exceeds the critical threshold of 95% in real time, the system will temporarily adjust the decision boundary of its security strategy based on a core cybernetics principle—that is, dynamically optimizing the objective function under multiple constraints. Specifically, this includes adjusting the decision function... Key threshold parameters A temporary increase of 0.05 (e.g., from 0.1 to 0.15) is implemented. This is not a reduction in the safety baseline, but rather a context-aware optimization. Essentially, when production pressure reaches its peak, the system intelligently and temporarily increases its tolerance for normal fluctuations, proactively avoiding triggering moderate-intensity defensive measures (such as end-to-end encryption) that, while increasing marginal safety gains, would severely compromise production continuity. This prevents security protection itself from becoming a bottleneck for production. This mechanism ensures that, under high load, the system prioritizes safeguarding absolute production stability while maintaining a minimum acceptable level of security.
[0093] Secondly, this embodiment integrates a lightweight game engine, specifically including hardware-accelerated parallel Nash equilibrium solving. Traditionally, solving for the Nash equilibrium of a game in software (especially in mixed-strategy scenarios involving dozens of nodes) is a computationally intensive and time-consuming process. In this embodiment, however, the core game payoff calculation and strategy iteration process is hardened from a general-purpose CPU instruction set to FPGA (Field-Programmable Gate Array) hardware logic circuits designed specifically for parallel computing. On the FPGA, payoff calculations for different device node pairs can be performed completely synchronously, and the computation process itself is designed as an efficient multi-stage pipeline. This fundamental shift from "serial software execution" to "parallel hardware execution" enables the system to solve a production line game model with 50 nodes within 8 milliseconds. This ultra-low latency makes it possible, for the first time in engineering practice, to guide proactive defense decisions based on real-time, accurate game inference, thereby elevating security response from a passive reaction to an intelligent prediction level.
[0094] Finally, in this embodiment, seamless execution with the OpenHarmony operating system was achieved, based on deterministic scheduling via a distributed soft bus and a microkernel. Traditional industrial network communication suffers from latency jitter, and policy deployment requires a lengthy protocol stack. OpenHarmony's distributed soft bus provides a unified virtual communication plane for all heterogeneous devices. After security policies are encapsulated as standard data objects, they can be directly routed with near-zero copying via the optimal path, avoiding multiple data copies within the kernel. Simultaneously, its microkernel architecture defines policy deployment tasks as the highest-priority real-time tasks, ensuring they are scheduled and executed immediately upon readiness. This principle upgrades policy deployment from "uncertain network message passing" to "deterministic system service calls," completely eliminating the uncertainty caused by protocol conversion and scheduling delays, achieving an end-to-end latency of less than 2 milliseconds from decision generation to implementation on the target device. This enables the system to perform highly granular operations such as "dynamically enabling encryption within a single control cycle to intercept malicious instructions in the next frame," achieving true tight coupling and synchronization between security control and production control.
[0095] In other words, the automatic policy degradation mechanism endows the system with a deep understanding and dynamic adaptability to production business needs; the lightweight game engine provides the instantaneous computing power necessary to support intelligent decision-making; and OpenHarmony's seamless execution bridges the "last millisecond" deterministic channel between the policy bits in the digital world and the control atoms in the physical world. These three elements work together to construct an industrial defense entity with inherent security intelligence, capable of understanding production, anticipating threats, and acting instantly.
[0096] In this embodiment, before dynamically defending the target production line based on the target defense strategy, the method further includes: constructing a closed-loop feedback mechanism, and dynamically updating the preset defense threshold through gradient optimization based on the actual security effect and production impact of historical defense strategies.
[0097] The step of dynamically updating the preset defense threshold through gradient optimization includes: dynamically updating the preset defense threshold based on a target update formula; the target update formula is: ;in, Indicates the first The preset defense threshold for the next iteration Indicates the first The preset defense threshold for the next iteration Indicates the learning rate. This represents the overall objective function that maximizes profits. This indicates the relationship between the overall objective function and the preset defense threshold. Find the partial derivative. This partial derivative indicates the preset defense threshold at the current point. How does the overall efficiency (safety-production) change when a tiny unit is added? If the derivative is positive, it means that increasing the threshold improves overall efficiency; if it is negative, the opposite is true.
[0098] The target update formula is the gradient descent formula. With a certain learning rate... Along the path that maximizes ( Safety score (production loss) comprehensive effectiveness.
[0099] Specifically, in the specific implementation process, by comparing different The system automatically identifies the optimal threshold for the current production environment based on the actual safety effects and production impact of the proposed strategy.
[0100] This enables continuous learning and optimization. By regularly updating equipment behavior baselines (to adapt to equipment aging), rolling back test strategies during maintenance periods, and generating performance reports, the entire defense system can adapt to new threats and production changes, and continuously iterate and improve.
[0101] Specifically, the performance parameters of industrial equipment drift slowly due to natural wear and tear and aging. To address this, after each production batch, the system automatically utilizes the operational data collected within that batch, employing an incremental learning algorithm to adjust the equipment behavior baseline (including the mean of the state vector). Covariance and cluster synchronization benchmark The system iterates and updates the data. This process ensures that the "normal" state defined by the system continuously tracks the actual physical state of the device, preventing performance degradation from being misjudged as a safety event due to outdated baselines, thereby maintaining the long-term accuracy of the anomaly detection model.
[0102] Secondly, this embodiment incorporates a periodic defense strategy effectiveness verification mechanism. Utilizing the static window of the production line during planned maintenance periods (such as weekends), the system conducts controlled rollback tests on the currently deployed proactive defense strategy in an isolated sandbox environment. By replaying historical attack traffic and abnormal scenarios, and comparing the interception effects, resource overhead, and potential impact on simulated production processes of different strategy versions (such as the current version versus historical versions), the system can objectively evaluate strategy effectiveness and identify defense blind spots or redundant configurations. This process is equivalent to providing the security system with regular "real-world drills," ensuring that its defense logic remains optimal when facing real threats.
[0103] Ultimately, all learning and validation data are systematically analyzed, and a "Strategy Effectiveness Analysis Report" is automatically generated. This report quantifies key indicators such as detection accuracy, false positive rate, production delays caused by defensive actions, and the trade-off between security gains and production losses. The core value of the report lies in providing data-driven decision-making support, clearly indicating the preset defense thresholds as described. Game weight coefficients The report outlines optimization directions for core parameters. Based on this report, the system can automatically perform parameter tuning, such as applying gradient update rules to adjust the preset defense threshold. Alternatively, the administrator can approve the iterative upgrade of the policy logic to complete the complete closed loop from data awareness to policy improvement.
[0104] Based on this, by embedding baseline updates, policy verification, and report-driven optimization into the natural cycle of production and maintenance, a fundamental shift in security capabilities from static configuration to dynamic growth has been achieved, enabling the protection system to adaptively respond to equipment aging, process changes, and threat evolution.
[0105] In this embodiment, an integrated OpenHarmony system is used to build a basic operating environment that supports its high real-time performance, high reliability, and adaptive security capabilities. Refer to Figure 6 for details; Figure 6 illustrates the industrial scenario verification process. Its core integration points are reflected in three levels: First, instantaneous synchronization of collaborative states is achieved through a distributed soft bus: As an operating system-level unified communication plane, the distributed soft bus provides a low-latency, deterministic data exchange channel for all devices on the production line (PLCs, robotic arms, sensors). This enables the state vectors of each device to... The global information required for game theory simulation can be synchronized across nodes in real time with a latency of less than 10ms, enabling the construction of an accurate collaborative feature matrix. It provides physical guarantees for global game calculations and solves the problem of inconsistent security state views caused by protocol heterogeneity and latency jitter in traditional industrial networks.
[0106] Secondly, the integrity and confidentiality of the defense strategy are ensured by relying on a secure and trusted execution environment: the system will use the optimal defense strategy calculated by the game optimization engine ( or strategy distribution Encryption, signing, and storage are performed within the hardware-level trusted execution environment (TEE) provided by OpenHarmony. Policies are distributed through a secure channel and executed after decryption and verification within the target device's TEE. This ensures that the entire decision tree and response instruction chain cannot be tampered with, stolen, or bypassed during deployment and execution, thereby placing core security logic under the protection of a solidified root of trust to resist malicious interference from within the system or from compromised nodes.
[0107] Adaptive scheduling of defense resources using elastic security services: OpenHarmony's elastic service framework allows security services to dynamically request and release system computing resources based on real-time threat levels (ΔS / ΔS values). For example, when... When the threshold is greater than 0.5, triggering the "Game Theory Defense" mode, the system will automatically request and exclusively allocate FPGA accelerator resources through the elastic service framework to ensure that complex game models are solved within 8ms; while at low threat levels ( When the threshold is less than or equal to 0.2, dedicated computing power is released, and logging tasks are performed only on low-power CPU cores. This dynamic matching mechanism between computing power and threats ensures ultimate real-time response in high-end conditions while also optimizing the overall energy efficiency and resource utilization of the system.
[0108] As can be seen, the device collaborative security game-theoretic defense method proposed in this embodiment for the Industrial Internet of Things environment firstly constructs a dynamic baseline based on behavior, performing three-dimensional modeling of device command frequency, communication message information entropy, and response delay standard deviation to form a unique device digital fingerprint, thereby overcoming the adaptability defects of static rule bases to operating condition fluctuations. Secondly, the core lies in the quantitative decision engine of the attack-defense game, which inversely transforms the Game-Theoretic Attack Tree (GTA) into a defense strategy tree. Its built-in Nash equilibrium solver can dynamically weigh security utility against production loss. For example, when the attack risk is high, even if the defense measures may lead to an 8.7% production loss, they will still be enforced; when the risk is low, only logs are recorded, with an impact of less than 0.1%, and the probability distribution of attack paths is predicted with an accuracy of over 92% through sandbox simulation. Furthermore, a hierarchical strategy self-tuning mechanism is designed, driven by the security offset index (ΔS) to achieve precise adaptive adjustment of defense strength. Compared with the traditional "one-size-fits-all" approach, it can reduce the production loss caused by security protection by an average of 14.5 times. Finally, the entire solution is supported by deep integration with OpenHarmony, enabling millisecond-level deployment of defense strategies via its distributed soft bus (e.g., reducing the synchronization latency of heterogeneous device strategies from 150ms to 8ms). Its security kernel ensures the immutability of the game-theoretic decision tree, and the entire defense strategy kernel occupies less than 1.2MB of memory, sufficient for deployment on resource-constrained edge terminals. This invention not only accurately addresses the urgent needs of current industrial security but also achieves a fundamental leap from static, passive, and isolated protection to dynamic, proactive, and collaborative intelligent defense at the architectural level.
[0109] In summary, this embodiment provides a device collaborative security game-theoretic defense method in an industrial IoT environment. It acquires target data, constructs a target multi-dimensional device state vector and a target collaborative feature matrix based on the target data, including the operational data of each device on the target production line. Then, it constructs a role-action space mapping, builds a target payoff matrix based on the role-action space mapping, and obtains the target path probability distribution of the system attacker based on the target payoff matrix. Next, it obtains target anomaly deviations and target synchronization decay factors based on the target multi-dimensional device state vector and the target collaborative feature matrix, and constructs a target security offset index formula based on the target anomaly deviations and target synchronization decay factors, where the target anomaly deviations and target synchronization decay factors are respectively the anomaly deviations and decay factors of the devices on the target production line. Finally, it obtains a target defense strategy based on the target security offset index formula and the target path probability distribution, and dynamically defends the target production line based on the target defense strategy. The device collaborative security game-theoretic defense method proposed in this embodiment solves the problems of high false alarm rates and insufficient defense against complex attacks caused by the lack of accurate modeling of production line-level device collaborative behavior in existing technologies. By constructing an intelligent defense system that integrates real-time behavioral profiling, game theory-based proactive decision-making, and closed-loop self-evolution capabilities, it dynamically balances security protection and production efficiency on a millisecond-level timescale. This fundamentally solves the core contradictions in industrial environments, such as the inability to defend against advanced threats, the misuse of security measures, and the unseen nature of coordinated attacks, and achieves a paradigm shift from static passive protection to dynamic intrinsic security.
[0110] It should be understood that although the steps in the flowcharts shown in the accompanying drawings are displayed sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0111] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink), DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0112] Example 2: Based on the above embodiments, the present invention also provides a device collaborative security game defense system in an industrial Internet of Things (IoT) environment, as shown in Figure 7. The device collaborative security game defense system in an industrial IoT environment includes: a construction module, used to acquire target data, and construct a target multi-dimensional device state vector and a target collaborative feature matrix based on the target data. The target data includes the operating data of each device on the target production line, as described in Example 1; a collaborative modeling module, used to construct a role-action space mapping, construct a target payoff matrix based on the role-action space mapping, and obtain the target path probability distribution of the system attacker based on the target payoff matrix, specifically... As described in Embodiment 1; the security quantification module is used to obtain the target abnormal deviation and the target synchronization attenuation factor based on the target multi-dimensional equipment state vector and the target collaborative feature matrix, and to construct the target security offset index formula based on the target abnormal deviation and the target synchronization attenuation factor, wherein the target abnormal deviation and the target synchronization attenuation factor are respectively the abnormal deviation and attenuation factor of the equipment on the target production line, as specifically described in Embodiment 1; the strategy acquisition module is used to obtain the target defense strategy based on the target security offset index formula and the target path probability distribution, and to perform dynamic defense on the target production line based on the target defense strategy, as specifically described in Embodiment 1.
[0113] Example 3: Based on the above embodiments, the present invention also provides a terminal, as shown in FIG8, which includes a processor 10 and a memory 20. FIG8 only shows some components of the terminal; however, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead.
[0114] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard drive or RAM. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard drive, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard. Furthermore, the memory 20 may include both internal and external storage units. The memory 20 is used to store application software and various types of data installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output.
[0115] In one embodiment, the memory 20 stores a device collaborative security game defense program 30 in an industrial Internet of Things (IoT) environment. This device collaborative security game defense program 30 in an industrial Internet of Things (IoT) environment can be executed by the processor 10, thereby realizing the device collaborative security game defense method in an industrial Internet of Things (IoT) environment of this application.
[0116] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other chip, used to run program code stored in the memory 20 or process data, such as executing the device collaborative security game defense method in the industrial Internet of Things environment.
[0117] In one embodiment, when the processor 10 executes the device collaborative security game defense program 30 in the industrial IoT environment stored in the memory 20, the following steps are implemented: The device collaborative security game defense method in the industrial IoT environment includes: acquiring target data, constructing a target multi-dimensional device state vector and a target collaborative feature matrix based on the target data, wherein the target data includes the operating data of each device on the target production line; constructing a role-action space mapping, constructing a target payoff matrix based on the role-action space mapping, and obtaining the target path probability distribution of the system attacker based on the target payoff matrix; obtaining target abnormal deviation and target synchronization decay factor based on the target multi-dimensional device state vector and the target collaborative feature matrix, constructing a target security offset index formula based on the target abnormal deviation and the target synchronization decay factor, wherein the target abnormal deviation and the target synchronization decay factor are respectively the abnormal deviation and decay factor of the devices on the target production line; obtaining a target defense strategy based on the target security offset index formula and the target path probability distribution, and performing dynamic defense on the target production line based on the target defense strategy.
[0118] Example 4: The present invention also provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement the steps of the device collaborative security game defense method in the industrial Internet of Things environment as described above.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A device collaborative security game defense method in an industrial Internet of Things environment, characterized in that, The device collaborative security game defense method in the industrial Internet of Things environment includes: acquiring target data; constructing a target multi-dimensional device state vector and a target collaborative feature matrix based on the target data, wherein the target data includes the operating data of each device on the target production line; constructing a role-action space mapping; constructing a target payoff matrix based on the role-action space mapping; and obtaining the target path probability distribution of the system attacker based on the target payoff matrix; obtaining target abnormal deviation and target synchronization decay factor based on the target multi-dimensional device state vector and the target collaborative feature matrix; constructing a target security offset index formula based on the target abnormal deviation and the target synchronization decay factor, wherein the target abnormal deviation and the target synchronization decay factor are respectively the abnormal deviation and decay factor of the devices on the target production line; obtaining a target defense strategy based on the target security offset index formula and the target path probability distribution; and performing dynamic defense on the target production line based on the target defense strategy.
2. The device collaborative security game defense method in the industrial Internet of Things environment according to claim 1, characterized in that, The target multidimensional device state vector is: ;in, Let be the state vector of device d at time t. Represents matrix transpose. This represents the PLC instruction sequence for device d. Indicates from A function to extract instruction execution frequency characteristics; Let d be the set of communication messages of device d. for A function for calculating information entropy; Let d be the response time series of device d. From A function to extract the delay fluctuation characteristics.
3. The device collaborative security game defense method in the industrial Internet of Things environment according to claim 1, characterized in that, The target collaborative feature matrix is: ;in, Indicates device and equipment Dependency relationship between them express and Statistical dependence between them For equipment State vector, For equipment State vector; This indicates the synchronization relationship of all equipment on the target production line. This represents the set of all equipment on the target production line. for The momentum, for The orientation angle in three-dimensional space, for The average value of the orientation angle of the state vectors of all devices in the system.
4. The device collaborative security game defense method in the industrial Internet of Things environment according to claim 1, characterized in that, The role-action space mapping is as follows: ;in, This provides the action space for system attackers. Let be the strategy probability vector of the attacker in the system; For the action space of the system defenders, Let be the system defender strategy probability vector.
5. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 4, characterized in that, The target return matrix is as follows: ; ; ;in, For the utility of the system defender, The weighting coefficients and , For the safety scoring function, This represents the current system state. The actions of the attacker in the system. For the actions of the system defenders, for Production losses; For the safety scoring formula, The current state vector With baseline state vector The Euclidean distance between them This is the maximum acceptable deviation distance; The formula for production loss rate is as follows: To execute The delay time, This is the initial cycle time.
6. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 5, characterized in that, The probability distribution of the target path is as follows: ; System status Under these conditions, the attacker selects an action in the system. The conditional probability; For a temperature parameter greater than 0, Indicates the attacker is in state Next, take action And the system defenders take the optimal response. At that time, the gains obtained by the attacker of the system, This represents any attack action by the attacker in the system; argmax is the parameter used to make the following function reach its maximum value.
7. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 6, characterized in that, The formula for the target safety offset index is: ; ; ;in, Indicates device At any moment The aforementioned target anomaly deviation, Indicates device The current state vector, Indicates device The mean vector of the historical normal state, Indicates device Covariance matrix of historical normal state vector The inverse matrix; This represents the target synchronization attenuation factor of the target production line. For the target production line at time The actual synchronization rate, This is the baseline synchronization rate of the target production line under historical normal operating conditions; Indicates all individual devices Find the average value. Indicates the target production line Find the average value.
8. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 4, characterized in that, The target defense strategy is obtained based on the target security offset index formula and the target path probability distribution, including: obtaining a preset defense threshold, which includes a first threshold and a second threshold, wherein the first threshold is used for conventional production lines and the second threshold is used for high-precision manufacturing production lines, and the second threshold is less than the first threshold; and obtaining the target defense strategy based on the preset defense threshold.
9. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 8, characterized in that, Obtaining the target defense strategy based on the preset defense threshold includes: constructing a hierarchical strategy decision tree based on the target security offset index formula and the preset defense threshold; obtaining the target risk level based on the hierarchical strategy decision tree, wherein the target risk level includes low risk, medium risk and high risk; and obtaining the target defense strategy based on the target risk level.
10. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 9, characterized in that, The hierarchical strategy decision tree is as follows: ;in, The preset defense threshold, Represents the policy decision function. The aforementioned low risk; The aforementioned medium risk; This is considered a high-risk situation.
11. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 9, characterized in that, The target defense strategy is obtained based on the target risk level, including: if the target risk level is low risk, the target defense strategy is to only log; if the target risk level is medium risk, the target defense strategy is encryption and authentication; if the target risk level is high risk, the target defense strategy is game theory active defense.
12. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 9, characterized in that, Before obtaining the target defense strategy, the method further includes: detecting the current capacity utilization rate of the target production line; if the current capacity utilization rate is higher than the target threshold, then increasing the preset defense threshold by 0.
05.
13. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 11, characterized in that, Obtaining the target defense strategy based on the target risk level includes: when the target risk level is high risk, obtaining a game-theoretic optimization formula, and optimizing the target defense strategy based on the game-theoretic optimization formula; the game-theoretic optimization formula is: ;in, This indicates that the predicted attack actions against the attacker targeting the system are taken at the maximum value. This indicates that the defensive action taken by the defender against the system is the minimum value. This indicates the gains of the attacker in the system. Benefits of the system defender The expected value of the difference For the actions of the system defender The impact on the production of the target production line is greater than or equal to 0.
14. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 13, characterized in that, The method for obtaining the target defense strategy based on the target risk level further includes: constructing a strategy deployment function based on the target path probability distribution, and implementing the game-theoretic active defense strategy based on the strategy deployment function; the strategy deployment function is: ;in, For activation function, For the secure kernel of the OpenHarmony operating system, This represents the probability distribution vector of the target defense strategy. Represents a timestamp.
15. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 9, characterized in that, Before dynamically defending the target production line based on the target defense strategy, the method further includes: constructing a closed-loop feedback mechanism, and dynamically updating the preset defense threshold through gradient optimization based on the actual security effect and production impact of historical defense strategies.
16. The device collaborative security game defense method in an industrial Internet of Things environment according to claim 15, characterized in that, The step of dynamically updating the preset defense threshold through gradient optimization includes: dynamically updating the preset defense threshold based on a target update formula; the target update formula is: ;in, Indicates the first The preset defense threshold for the next iteration Indicates the first The preset defense threshold for the next iteration Indicates the learning rate. This represents the overall objective function that maximizes profits. Indicates to Find the partial derivatives.
17. A device collaborative security game defense system in an industrial Internet of Things environment, characterized in that, include: A construction module is used to acquire target data and construct a target multidimensional equipment state vector and a target collaborative feature matrix based on the target data. The target data includes the operating data of each device on the target production line. The collaborative modeling module is used to construct a role-action space mapping, build a target reward matrix based on the role-action space mapping, and obtain the target path probability distribution of the system attacker based on the target reward matrix; The safety quantification module is used to obtain the target abnormal deviation and the target synchronization attenuation factor based on the target multi-dimensional equipment state vector and the target collaborative feature matrix, and to construct the target safety offset index formula based on the target abnormal deviation and the target synchronization attenuation factor, wherein the target abnormal deviation and the target synchronization attenuation factor are respectively the abnormal deviation and attenuation factor of the equipment on the target production line; The strategy acquisition module is used to acquire a target defense strategy based on the target security offset index formula and the target path probability distribution, and to perform dynamic defense on the target production line based on the target defense strategy.
18. The device collaborative security game defense system in an industrial Internet of Things environment according to claim 17, characterized in that, Also includes: The feedback mechanism construction module is used to build a closed-loop feedback mechanism, which dynamically updates the preset defense threshold through gradient optimization based on the actual security effect and production impact of historical defense strategies.
19. A terminal, characterized in that, The terminal includes: a processor and a computer-readable storage medium communicatively connected to the processor, the computer-readable storage medium being adapted to store multiple instructions, and the processor being adapted to invoke the instructions in the computer-readable storage medium to execute the steps of implementing the device collaborative security game defense method in the industrial Internet of Things environment as described in any one of claims 1-7.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the device collaborative security game defense method in an industrial Internet of Things environment as described in any one of claims 1-7.