An industrial control system dynamic protection method based on TEE and reinforcement learning

CN122783274APending Publication Date: 2026-09-18INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610809307.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

现有基于人工智能的安全方案虽具备一定的行为分析与预测能力,但其模型通常运行于普通执行环境(REE)中,易受到恶意软件篡改、数据窃取或边信道攻击,导致模型失效或输出被操控,严重削弱了系统的可信度

Benefits of technology

本申请通过激活硬件信任根并划分TEE/REE隔离内存空间,在硬件层面构建安全飞地,确保安全基线数据、加密密钥、强化学习模型等核心资产仅在受保护环境中加载与运行,从根本上防止外部恶意程序的窥探与篡改,显著提升了AI驱动安全机制的可信等级。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122783274A_ABST
    Figure CN122783274A_ABST
Patent Text Reader

Abstract

The application provides a kind of industrial control system dynamic protection method based on TEE and reinforcement learning, it is related to industrial control system information security technical field, including: activating hardware trust root, the isolated memory space of trusted execution environment is divided according to hardware trust root, and presetting security baseline data in the isolated memory space of trusted execution environment;In the isolated memory space of trusted execution environment, the multi-source data is encrypted and collected, the credibility of the source node of multi-source data is verified by remote proof, while the logical consistency between multi-source data is analyzed to identify anomalies, the network traffic features, device behavior features and environmental state features are included;Network traffic features, device behavior features and environmental state features are integrated into state vector, input into reinforcement learning model running in the isolated memory space of trusted execution environment for online inference, output optimal security action and intercept abnormal instruction in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of information security technology for industrial control systems, specifically relating to a dynamic protection method for industrial control systems based on TEE and reinforcement learning. Background Technology

[0002] With the rapid development of the Industrial Internet, industrial control systems are gradually shifting from closed to open, adopting a large number of common protocols, hardware and software platforms, and network interconnections, leading to increasingly severe cybersecurity threats. Traditional industrial control system security protection mainly relies on static rule-based or signature-matching technologies such as firewalls and intrusion detection systems (IDS), which are insufficient to cope with new and complex attacks such as advanced persistent threats (APTs), zero-day vulnerability attacks, and command tampering. These attacks are often characterized by strong concealment, long latency periods, and dynamically changing behavioral patterns. Traditional methods, lacking environmental awareness and adaptive response capabilities, cannot effectively identify and block attacks in their early stages.

[0003] Furthermore, industrial control systems have extremely high real-time requirements, with control command execution latency typically needing to be controlled within milliseconds or even microseconds. While existing AI-based security solutions possess some behavioral analysis and prediction capabilities, their models usually run in a Free Execution Environment (REE), making them vulnerable to malware tampering, data theft, or side-channel attacks. This can lead to model failure or output manipulation, severely undermining system reliability. Simultaneously, AI models themselves have significant computational overhead, easily causing performance bottlenecks when deployed on resource-constrained edge devices (such as PLCs and sensor gateways), impacting production process stability.

[0004] While Trusted Execution Environment (TEE) technology can protect sensitive data and code through hardware-level isolation mechanisms (such as ARM TrustZone and Intel SGX), its applications are mostly limited to static encrypted storage and authentication, lacking proactive defense and intelligent decision-making capabilities, and unable to dynamically adjust security policies based on network and device status. Therefore, how to achieve real-time, intelligent, and adaptive security protection for industrial control systems while ensuring the safe operation of AI models has become a pressing technical challenge. Summary of the Invention

[0005] This application provides a dynamic protection method for industrial control systems based on TEE and reinforcement learning to solve one of the aforementioned technical problems.

[0006] The technical solution adopted in this application is as follows: This application provides a dynamic protection method for industrial control systems based on TEE and reinforcement learning, including: Activate the hardware root of trust, divide the isolated memory space of the trusted execution environment according to the hardware root of trust, and preset security baseline data in the isolated memory space of the trusted execution environment. Encrypted collection of multi-source data is performed in the isolated memory space of the trusted execution environment. The trustworthiness of the source nodes of the multi-source data is verified by remote proof. At the same time, the logical consistency between the multi-source data is analyzed to identify anomalies. The multi-source data includes network traffic characteristics, device behavior characteristics and environmental status characteristics. Network traffic characteristics, device behavior characteristics, and environmental state characteristics are integrated into a state vector, which is then input into a reinforcement learning model running in the isolated memory space of a trusted execution environment for online inference, outputting the optimal security action and intercepting abnormal commands in real time.

[0007] According to one embodiment of this application, it also includes: The reward function is dynamically adjusted based on the feedback results of safety actions, the parameters of the reinforcement learning model are locally updated, and the encrypted gradient parameters of multiple plant areas are aggregated through a federated learning mechanism to generate a global model. According to one embodiment of this application, it also includes: When configuration file tampering is detected, the device state is restored by calling the security baseline data stored in the isolated memory space of the trusted execution environment, and the operation log is signed and stored on the blockchain to form a closed-loop defense ecosystem.

[0008] According to one embodiment of this application, activating the hardware root of trust involves dividing the trusted execution environment into isolated memory spaces based on the hardware root of trust, and pre-setting security baseline data in the isolated memory spaces of the trusted execution environment, specifically as follows: Activate the hardware root of trust and generate a unique hardware identifier for the device through a physically unclonable function for identity authentication and integrity verification; A hardware acceleration module for the SM4 / SM9 national cryptographic algorithm is loaded within the isolated memory space of the trusted execution environment to achieve real-time encrypted transmission of control commands and sensor data. The digital signature within the isolated memory space of the trusted execution environment is verified through a hardware root of trust to ensure it has not been tampered with.

[0009] According to one embodiment of this application, multi-source data is encrypted and collected in an isolated memory space of a trusted execution environment. The trustworthiness of the source nodes of the multi-source data is verified through remote proof, and the logical consistency between the multi-source data is analyzed to identify anomalies. The multi-source data includes network traffic characteristics, device behavior characteristics, and environmental state characteristics, specifically: Real-time capture of communication data from industrial protocols is used to encrypt and transmit network traffic characteristics to an isolated memory space within a trusted execution environment. Data from temperature, vibration, and current sensors are collected and synchronized with control commands to form a time-series data set as a characteristic of the equipment's behavior. Remote verification verifies whether the data source node is running in the isolated memory space of a trusted execution environment, preventing counterfeit devices from accessing the site. The logical relationships between multi-sensor data are analyzed within the isolated memory space of the trusted execution environment as environmental state characteristics. When the deviation exceeds a preset threshold, it is marked as suspicious data and an alarm is triggered.

[0010] According to one embodiment of this application, the process of integrating network traffic characteristics, device behavior characteristics, and environmental state characteristics into a state vector, inputting it into a reinforcement learning model running in the isolated memory space of a trusted execution environment for online inference, outputting the optimal security action, and intercepting abnormal instructions in real time, specifically involves: Construct a state vector that includes network traffic characteristics, device behavior characteristics, and environmental state characteristics; A pruned and optimized deep Q-network model is run in the isolated memory space of a trusted execution environment. Forward inference is performed based on the state vector, and the expected value score of each security action is output. Select the highest-rated safety action, including logging, issuing alarms, isolating the device, or switching redundant controllers; When a PLC instruction is detected to exceed the baseline safety data by ±20%, the abnormal instruction is immediately intercepted and its transmission to the actuator is blocked.

[0011] According to one embodiment of this application, the step of dynamically adjusting the reward function based on the feedback results of safety actions, locally updating the reinforcement learning model parameters, and aggregating the encryption gradient parameters of multiple plant areas through a federated learning mechanism to generate a global model specifically involves: Rewards are calculated based on the results of safety actions. Successful interception earns a positive reward, while false alarms or missed alarms incur negative penalties. Incremental training of the Q-network parameters of the local DQN model is performed using reward signals within the isolated memory space of the trusted execution environment, thereby improving the ability to identify attack patterns specific to the plant area. The updated gradient parameters of the local model are encrypted using an SM9 public key and a signature is attached to the isolated memory space of the trusted execution environment before being uploaded to the central server. The central server decrypts and aggregates the gradients of each node in the isolated memory space of the trusted execution environment, generates an updated global model, and encrypts and distributes it to each edge node. After verifying the global model signature, the edge nodes replace the local model parameters, thus completing the periodic evolution of the policy library.

[0012] According to one embodiment of this application, when configuration file tampering is detected, the device state is restored by calling the security baseline data stored in the isolated memory space of the trusted execution environment, and the operation log is signed and uploaded to the blockchain for evidence storage, forming a closed-loop defense ecosystem. Specifically: Continuously monitor the hash values ​​of PLC configuration programs and SCADA project files. When a discrepancy is found between the hash value and the security baseline stored in the isolated memory space of the trusted execution environment, it is determined that the file has been tampered with. The original configuration file, which is encrypted and stored in the isolated memory space of the trusted execution environment, is automatically recalled and re-burned to the PLC to restore it to a trusted operating state. The operation logs of all security events are signed with a private key in the isolated memory space of the trusted execution environment, and the signed logs are then uploaded to the blockchain.

[0013] A second aspect of this application provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps described in the method.

[0014] A third aspect of this application provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method as described.

[0015] Due to the adoption of the above technical solution, the beneficial effects achieved by this application are as follows: This application constructs a secure enclave at the hardware level by activating the hardware root of trust and dividing the TEE / REE isolated memory space. This ensures that core assets such as security baseline data, encryption keys, and reinforcement learning models are loaded and run only in the protected environment, fundamentally preventing external malicious programs from spying on and tampering with them, and significantly improving the trust level of the AI-driven security mechanism.

[0016] This application encrypts and collects multi-source data such as network traffic, device behavior, and environmental status in a TEE environment, and verifies the identity and credibility of the data source node by combining a remote authentication mechanism, effectively resisting the risk of counterfeit device access; at the same time, by analyzing the logical consistency between multi-sensor data (such as the coupling relationship between vibration and temperature), it can identify sensor data tampering or deception attacks, reduce false alarm rate and false negative rate, and improve the accuracy of threat detection.

[0017] This application integrates multi-dimensional state features into a unified state vector, which is then input into a reinforcement learning model running within the TEE for online inference. This enables millisecond-level identification and real-time interception of abnormal commands (such as PLC commands exceeding the process baseline by ±20%). This mechanism overcomes the limitations of traditional static rules, enabling the autonomous generation of optimal response strategies under unknown attack modes, significantly enhancing the system's dynamic defense capabilities.

[0018] This application's reinforcement learning model continuously receives feedback from defensive actions within the TEE (Time-of-Effect), dynamically adjusts the reward function, and optimizes the policy output, forming a closed-loop mechanism of "perception-decision-execution-learning." This design enables the security policy to have self-evolution capabilities, automatically updating as attack patterns evolve, avoiding the lag caused by manual intervention, and significantly improving the long-term effectiveness of the protection system.

[0019] All security processing steps in this application are completed within a TEE (Trusted Edge Environment). By combining hardware acceleration of national cryptographic algorithms with a lightweight model design, the end-to-end latency from data acquisition to policy execution is ensured to meet the stringent real-time requirements of industrial control systems. Furthermore, since core computations are concentrated in the edge trusted environment, there is no need to frequently upload raw data to the cloud, reducing network bandwidth consumption and centralized processing pressure. Attached Figure Description

[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a dynamic protection method for industrial control systems based on TEE and reinforcement learning, provided for an embodiment of this application; Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0021] Figure label: 810, Processor; 820, Communication interface; 830, Memory; 840, Communication bus. Detailed Implementation

[0022] To more clearly illustrate the overall concept of this application, a detailed explanation is provided below with reference to the accompanying drawings.

[0023] Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below. It should be noted that, unless otherwise specified, the embodiments of this application and the features thereof can be combined with each other.

[0024] In this application, unless otherwise expressly specified and limited, the "above" or "below" of the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. In the description of this specification, references to terms such as "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples.

[0025] Example 1 like Figure 1 As shown, a dynamic protection method for industrial control systems based on TEE and reinforcement learning includes: Activate the hardware root of trust, divide the isolated memory space of the trusted execution environment according to the hardware root of trust, and preset security baseline data in the isolated memory space of the trusted execution environment.

[0026] As mentioned above, firstly, "activating the hardware root of trust" refers to generating a unique, uncopyable identity for the device through an immutable hardware module built into the chip (such as a Physically Unclonable Function (PUF) or a secure boot engine) when the system powers on or restarts. This identity serves as the starting point for the trustworthiness of the entire system. This root of trust has anti-physical extraction and anti-cloning properties, and is the foundation for building a trusted computing system.

[0027] Subsequently, "isolating the Trusted Execution Environment (TEE) memory space based on the hardware root of trust" refers to using hardware-level memory protection mechanisms (such as memory encryption, access control, address space isolation, etc.) to create a secure area in system memory that is completely isolated from the ordinary execution environment (REE). This area only allows trusted code and data verified by digital signatures to be loaded and executed. External operating systems or malicious programs cannot read or modify its contents, thereby achieving strong protection for sensitive information and critical logic.

[0028] Finally, "pre-setting secure baseline data in the isolated memory space of the Trusted Execution Environment" refers to pre-storing critical reference information required for the normal operation of the industrial control system (such as PLC instruction cycle range, set of valid opcodes for control instructions, normal value ranges of sensors, and process flow timing logic) in a trusted storage area within the TEE in an encrypted manner. This baseline data serves as the basis for comparison in subsequent anomaly detection and strategy decision-making. Its integrity and confidentiality are guaranteed by the TEE, preventing attackers from tampering with or stealing it.

[0029] For example, let's take the wind turbine control system of a wind farm as an example: When the wind turbine controller (such as a PLC or edge gateway) is powered on, it first activates the hardware root of trust module of its built-in RISC-V chip. This module generates a globally unique device fingerprint through a Physically Unclonable Function (PUF) and uses it for subsequent identity authentication.

[0030] Based on this root of trust, the system invokes the chip's security extension functions (such as TrustZone technology) to allocate an independent, hardware-protected TEE region in memory. During this process, the system performs hash verification on the TEE kernel image; only legitimate images that pass signature verification can be loaded and run, preventing malicious firmware injection.

[0031] Subsequently, the system loads pre-configured security baseline data into the encrypted storage area within the TEE. For example, under normal operating conditions, the wind turbine should send pitch angle adjustment commands every 500 milliseconds, with an allowable deviation of ±10%; temperature sensor readings should be between -20℃ and 85℃; and the dominant frequency of the vibration spectrum should be concentrated in the range of 15Hz to 25Hz. All this data is stored in encrypted form within the TEE and is only decrypted and used by trusted code when comparison is required.

[0032] Subsequently, all security-related operations, such as command legitimacy judgment, abnormal behavior identification, and model reasoning decision-making, are executed in this TEE environment to ensure that the entire protection link is in a trusted state from the beginning.

[0033] It should be noted that, in specific implementation scenarios, a master-slave trust root architecture can be set up based on the above solution. The hardware trust root of the master device (such as the central control server) is used to verify the trust status of multiple edge nodes, so as to achieve unified and trusted management across factories and networks.

[0034] In specific implementation scenarios, based on the above solutions, the security baseline data is not limited to static presets, but can also support receiving remote update instructions from the security management center through a trusted channel within the TEE, thereby achieving controlled upgrades of the baseline version to adapt to process changes or equipment replacement scenarios.

[0035] In specific implementation scenarios, based on the above solutions, security baselines can be classified by type (such as communication protocol baseline, device behavior baseline, physical environment baseline) and different access permission levels can be set within the TEE to ensure that highly sensitive data is only accessed by specific trusted modules.

[0036] In specific implementation scenarios, based on the above solution, the security baseline data can be digested using the SM3 hash algorithm within the TEE and protected using SM2 digital signatures. Any unauthorized modification to the baseline can be detected and alerted immediately.

[0037] In specific implementation scenarios, based on the above solutions, this trusted startup mechanism is not only applicable to wind power control systems, but can also be extended to other industrial scenarios such as intelligent manufacturing production lines, rail transit signaling systems, and chemical process control. The overall architecture can be reused simply by presetting the corresponding safety baseline according to different process characteristics.

[0038] Encrypted collection of multi-source data is performed in the isolated memory space of the trusted execution environment. The trustworthiness of the source nodes of the multi-source data is verified by remote proof. At the same time, the logical consistency between the multi-source data is analyzed to identify anomalies. The multi-source data includes network traffic characteristics, device behavior characteristics, and environmental state characteristics.

[0039] As mentioned above, firstly, "encrypting and collecting multi-source data in the isolated memory space of the Trusted Execution Environment" means that all data from industrial networks and physical devices is immediately decrypted and processed within the TEE after entering the system, preventing sensitive information from being exposed to the ordinary execution environment. The collected "multi-source data" covers three core types of information: network traffic characteristics (such as Modbus / TCP protocol communication frequency, message length distribution, and port access patterns), device behavior characteristics (such as PLC instruction execution cycle, SCADA operation sequence, and controller state switching frequency), and environmental status characteristics (such as sensor readings for temperature, vibration, current, and pressure). This data is captured in real time by the industrial protocol parsing module and encrypted for transmission and storage within the TEE using the national cryptographic SM4 algorithm to prevent man-in-the-middle theft or tampering. Secondly, "verifying the trustworthiness of the source node for multi-source data through remote verification" refers to the process where, before data reception, the local TEE initiates a remote verification request to the data sender (such as a remote PLC, sensor gateway, or edge controller). The sender needs to generate a verification report containing a hash value of its operating status within its own TEE environment and sign it using its private key. The receiving TEE verifies whether the signature and hash value match a pre-registered trusted image, confirming that the recipient's device has not been implanted with malicious firmware or subjected to jailbreak attacks, thereby ensuring the authenticity of the data source and the security of the device's operating environment.

[0040] Finally, "analyzing the logical consistency between multi-source data to identify anomalies" refers to performing correlation verification on the collected multi-dimensional data within the TEE. For example, when the motor load increases, its current should rise synchronously, and the vibration amplitude should increase. If phenomena such as "high current but no vibration" or "frequent commands but unchanged temperature" occur, which violate physical laws, they are judged as potential attacks or sensor spoofing. This mechanism does not rely on a single indicator threshold but establishes an expected relationship model between multi-source data based on the process mechanism, significantly improving the ability to identify covert attacks.

[0041] For example, consider the control system of a chemical reactor: During the reaction, the PLC sends a heating control command every 10 seconds, while the temperature sensor reports the temperature value every 5 seconds, and the current monitoring module of the stirring motor provides real-time feedback on the load status. Under normal circumstances, after the heating command is initiated, the temperature should rise exponentially, with slight fluctuations in the motor current. In this solution, all the aforementioned data is transmitted to the TEE environment of the central security gateway via an encrypted channel. Before receiving each piece of temperature data, the gateway initiates a remote authentication request to the sensor node to confirm that it is running in a trusted firmware environment and has not been hijacked. If the authentication fails, the node's data is discarded and an alarm is triggered.

[0042] Subsequently, the security analysis module within the TEE comprehensively compares the current network traffic (such as command sending frequency), device behavior (such as whether the PLC sends non-periodic commands), and environmental conditions (temperature change rate, current value). If the system detects that the PLC has not sent a heating command, but the temperature sensor data shows an abnormally high temperature and no change in motor current, which violates the physical logic that "heating requires power and stirring requires a load," the system determines that the sensor data has been forged or subjected to a man-in-the-middle injection attack, immediately triggering an alarm and suspending the relevant control processes.

[0043] It should be noted that, in specific implementation scenarios, the above solution can be extended to support data parsing and encryption encapsulation of mainstream industrial communication protocols such as PROFINET, OPC UA, and CANopen, enabling unified and reliable data acquisition across vendor devices.

[0044] In specific implementation scenarios, in addition to the above solutions, a lightweight causal reasoning model can be introduced into the TEE to automatically learn the dynamic correlation patterns between multiple source parameters based on historical operating data, and adapt to normal deviations caused by process parameter adjustments or equipment aging.

[0045] In specific implementation scenarios, based on the above solutions, a hierarchical remote verification structure can be constructed for large-scale distributed systems. Edge nodes prove their credibility to the upper-level gateway, and the gateway then proves the integrity of the entire link to the central security management platform, forming a defense-in-depth system.

[0046] In specific implementation scenarios, based on the above solutions and logical consistency analysis, a confidence assessment module can be introduced to generate anomaly scores based on factors such as the degree of deviation, duration, and scope of impact. These scores can then be used as input features for subsequent reinforcement learning models to improve decision-making accuracy.

[0047] In specific implementation scenarios, based on the above solutions, for maintenance terminals or mobile inspection equipment that need to temporarily access the system, a temporary trusted channel can be established within the TEE through a one-time remote certificate + time-limited access token, balancing security and operational flexibility.

[0048] In specific implementation scenarios, based on the above solutions, a high-precision timestamp synchronization function (such as the IEEE 1588 protocol) can be integrated into the TEE to ensure time alignment of multi-source data and avoid logical misjudgments caused by clock deviations.

[0049] In specific implementation scenarios, based on the above solutions, the identified abnormal data combinations (such as "no instruction + high temperature + low current") can be automatically generated into feature tags and encrypted and uploaded to the federated learning center for global model training, thereby improving the collaborative defense capabilities of multiple plant areas.

[0050] Network traffic characteristics, device behavior characteristics, and environmental state characteristics are integrated into a state vector, which is then input into a reinforcement learning model running in the isolated memory space of a trusted execution environment for online inference, outputting the optimal security action and intercepting abnormal commands in real time.

[0051] As mentioned above, firstly, "integrating network traffic characteristics, device behavior characteristics, and environmental state characteristics into a state vector" refers to normalizing and structuring real-time monitoring data from different dimensions within the TEE to form a unified input representation that can be understood by the model. Specifically: Network traffic characteristics include communication frequency, message length distribution, protocol compliance, and port access patterns; Equipment behavior characteristics include PLC instruction cycle stability, control command sequence regularity, and SCADA operation frequency; Environmental condition characteristics include readings from physical sensors such as temperature, vibration, current, and pressure, as well as their changing trends.

[0052] After standardization, these features are combined in a preset order to form a multidimensional numerical vector, which serves as the input to the current "environment state" of the reinforcement learning model.

[0053] Subsequently, "the input to the reinforcement learning model running in the isolated memory space of the Trusted Execution Environment (TEE) for online inference" refers to the state vector being fed into a lightweight reinforcement learning model (such as a pruned and optimized deep Q-network) deployed within the TEE. The model evaluates the value of each possible safe action based on the current policy and selects the action with the highest expected reward as the output. Because the model runs under hardware-level isolation protection throughout, its parameters and inference process are not subject to external interference, ensuring the integrity and confidentiality of the decision-making process.

[0054] Finally, "outputting optimal safety actions and intercepting abnormal commands in real time" means that the action commands output by the model (such as "logging", "issuing alarms", "blocking communication", or "switching to backup controller") are immediately parsed and executed by the safety execution module within the TEE. For detected abnormal control commands (such as non-periodic injection, illegal opcodes, parameter settings exceeding the process baseline by ±20%), the system can complete identification and interception within milliseconds to prevent them from being transmitted to the actuators and causing equipment malfunctions or production accidents.

[0055] For example, let's take a CNC machine tool processing production line as an example: During normal machine tool operation, the PLC sends a feed command every 100 milliseconds, the spindle speed remains at 3000±100 rpm, the coolant flow rate is stable, and there are no abnormal scanning behaviors in network communication. At a certain moment, the system monitors the following status: Network traffic characteristics: High-frequency probing behavior was observed at the PLC port; Equipment behavior characteristics: The PLC receives a non-periodic "emergency stop release + high-speed feed" compound instruction; Environmental characteristics: The spindle current suddenly increases, but the vibration does not increase synchronously, which violates the normal load response law.

[0056] After the above three types of features are collected, they are integrated into a single state vector within the TEE and input into the reinforcement learning model running within it. Based on historical training experience, the model determines that this combined feature is highly likely to be an instruction injection attack launched by an attacker exploiting a zero-day vulnerability. If executed, it could lead to tool breakage or workpiece ejection.

[0057] The model then outputs the optimal safety action: "Immediately intercept the instruction and isolate the PLC communication channel." Upon receiving this decision, the safety control module within the TEE immediately blocks the transmission path of the abnormal instruction, severs the communication connection between the PLC and the host computer, and triggers an audible and visual alarm to notify maintenance personnel. The entire process takes less than 80 milliseconds, far below the real-time threshold of the machine tool control system, effectively preventing safety accidents.

[0058] It should be noted that, in specific implementation scenarios, a model running container can be designed within the TEE based on the above solution, supporting the dynamic loading and switching of different types of reinforcement learning models such as DQN, PPO, and A3C. The optimal algorithm can be flexibly selected according to different application scenarios (such as choosing DQN for high real-time scenarios and PPO for high-complexity decision scenarios).

[0059] In specific implementation scenarios, a lightweight attention module can be added to the model input layer based on the above solution, enabling the model to automatically identify the most critical feature dimensions in the current state (such as network probing in the early stage of an attack and command anomalies in the later stage), thereby improving decision-making accuracy.

[0060] In specific implementation scenarios, the model output space can be expanded based on the above solutions to not only consider the attack blocking effect, but also take into account the impact on business continuity (such as prioritizing the degraded operation mode that does not affect production), thereby achieving collaborative optimization of security and production.

[0061] In specific implementation scenarios, based on the above solutions, the set of optional safety actions can be automatically adjusted according to the system's operating phase (such as startup, steady-state operation, and downtime maintenance). For example, more lenient instruction deviations can be allowed during the equipment startup phase, while sensitivity can be increased during high-load processing phases.

[0062] In specific implementation scenarios, a small-sample incremental update mechanism can be integrated into the TEE based on the above solution. When a new attack pattern is detected and confirmed by humans, the sample can be used to fine-tune the local model, improving the ability to respond quickly to similar attacks without waiting for a global model update.

[0063] In specific implementation scenarios, based on the above solution, before performing high-risk actions (such as device isolation or controller switching), the policy verification module embedded in the TEE can perform a secondary verification to confirm whether the action complies with preset security rules, preventing model misjudgment from leading to erroneous operations.

[0064] In specific implementation scenarios, based on the above solutions, the state vector can be synchronously mapped to a factory-level digital twin platform to simulate the impact of safety actions in a virtual environment, providing auxiliary reference for decision-making in complex scenarios.

[0065] In specific implementation scenarios, based on the above solution, for nodes with extremely limited computing resources, feature extraction can be completed within the TEE, and only the compressed state vector can be uploaded to the cloud trusted environment for auxiliary reasoning. The returned results can be verified and used for local decision-making, thus achieving a balance between resources and performance.

[0066] According to one embodiment of this application, it also includes: The reward function is dynamically adjusted based on the feedback results of safety actions, the parameters of the reinforcement learning model are locally updated, and the encrypted gradient parameters of multiple plant areas are aggregated through a federated learning mechanism to generate a global model. As described above, after each security action is executed, the system collects the actual execution result as feedback information, including whether the attack was successfully intercepted, whether false alarms or missed alarms occurred, and whether the operating status of the protected device was affected. This feedback information is sent to the Trusted Execution Environment (TEE) to dynamically adjust the reward function used by the reinforcement learning model. For example, when an isolation operation successfully prevents a confirmed attack, the system provides a positive reward; if a misjudgment causes the normal control flow to be interrupted, a negative penalty is imposed. In this way, the reward function can be continuously optimized based on the actual protection effect, guiding the model to learn more accurate decision-making strategies.

[0067] Based on the reward function adjustment, the locally running reinforcement learning model uses the feedback data to locally update its parameters. This process is completed within the TEE (Training Equipment Environment), ensuring that both the model training data and parameter changes are under hardware-level protection, preventing interference with the training process or malicious manipulation of the model. The updated model can adapt more quickly to the unique equipment behavior patterns and attack characteristics of this plant area, improving the accuracy of local protection.

[0068] To further enhance the model's generalization ability, the system employs a federated learning mechanism to achieve knowledge sharing among multiple plant areas. After each plant area completes its model parameter update locally, it does not upload the original data or the complete model. Instead, it extracts the gradient parameters generated by the update, encrypts them using the national cryptographic algorithm SM9, and attaches a digital signature generated by TEE to prove the node's trustworthiness. The encrypted gradient parameters are then uploaded to the central server.

[0069] After receiving encrypted gradient parameters from multiple plant areas, the central server first verifies the remote authentication information and signature validity of each node to confirm that it originates from legitimate and trusted device nodes. Only the verified gradient parameters are included in the aggregation process. Subsequently, these encrypted gradients are weighted and fused in a trusted environment to generate an updated global reinforcement learning model.

[0070] This global model integrates attack and defense experience from multiple plant areas, and is particularly adept at identifying new attack patterns that spread across regions (such as worm viruses targeting specific PLC models). The updated global model is encrypted and distributed to edge nodes in each plant area. Each node verifies the model signature within its TEE and completes the local model replacement, thereby achieving a synergistic improvement in overall protection capabilities. The entire process does not require sharing original operational data, ensuring data privacy while enabling continuous evolution of security strategies.

[0071] According to one embodiment of this application, it also includes: When configuration file tampering is detected, the device state is restored by calling the security baseline data stored in the isolated memory space of the trusted execution environment, and the operation log is signed and stored on the blockchain to form a closed-loop defense ecosystem.

[0072] As described above, the system continuously monitors the configuration files of industrial control equipment, including key operating parameters such as PLC program configuration, HMI screen settings, and SCADA project files. The monitoring process calculates the hash value of the file and compares it with secure baseline data stored in the isolated memory space of the Trusted Execution Environment (TEE). When the current file hash value is found to be inconsistent with the baseline, it is determined that the configuration file has been illegally tampered with, possibly due to an attacker injecting malicious code or modifying the control logic.

[0073] Upon confirmation of tampering, the system immediately triggers the recovery mechanism. The recovery operation is initiated within the TEE (Tracking Equipment), calling a pre-stored version of the security baseline—an encrypted copy of the trusted configuration file used when the device was operating normally. The system decrypts this security baseline data and writes it to the target device (such as a PLC or controller), restoring it to a known secure operating state. The entire recovery process is controlled by the TEE, ensuring the authenticity of the recovery source and the integrity of the recovery operation, preventing further interference from attackers during the recovery process.

[0074] Simultaneously, the system generates an operation log for this security incident, including the tamper detection time, information about the tampered device, original and current file hash values, and the execution status of recovery actions. This log is digitally signed within the Trusted Execution Environment (TEE) using the private key bound to the TEE, ensuring that the log content is unforgeable and unrepudiable.

[0075] Signed operation logs are uploaded to a blockchain-based evidence storage system (such as the FISCO BCOS consortium blockchain) via a secure interface, ensuring tamper-proof storage. The distributed nature of blockchain guarantees the long-term traceability of logs; even if the local system is compromised, audit records can still be verified on the chain. This mechanism meets the national graded protection system's requirement for secure log retention of no less than six months.

[0076] Through a complete process of "detection-recovery-recording-evidence storage," the system not only achieves rapid remediation of attack consequences but also constructs an auditable and traceable security closed loop. Operations personnel can perform post-incident analysis based on on-chain logs, identify attack paths, and feed typical event characteristics back to the reinforcement learning model for optimizing subsequent threat identification capabilities, thereby forming a continuously evolving closed-loop defense ecosystem.

[0077] According to one embodiment of this application, activating the hardware root of trust involves dividing the trusted execution environment into isolated memory spaces based on the hardware root of trust, and pre-setting security baseline data in the isolated memory spaces of the trusted execution environment, specifically as follows: Activate the hardware root of trust and generate a unique hardware identifier for the device through a physically unclonable function for identity authentication and integrity verification; A hardware acceleration module for the SM4 / SM9 national cryptographic algorithm is loaded within the isolated memory space of the trusted execution environment to achieve real-time encrypted transmission of control commands and sensor data. The digital signature within the isolated memory space of the trusted execution environment is verified through a hardware root of trust to ensure it has not been tampered with.

[0078] As described above, firstly, activating the hardware root of trust refers to the process of establishing a trust chain during the system startup phase through the trusted computing hardware module built into the chip. This hardware root of trust serves as the starting point for the entire system's security and possesses the characteristics of immutability and uniqueness. Based on this, a device-specific hardware identifier is generated using Physically Unclonable Function (PUF) technology. This identifier is determined by microscopic physical differences in the chip manufacturing process and cannot be copied or predicted; even on other devices of the same model, the same value cannot be generated. This unique hardware identifier is used in subsequent identity authentication processes to ensure that each device accessing the system possesses identifiable and unforgeable credentials, while also verifying the integrity of critical data and code to prevent unauthorized replacement or tampering.

[0079] Secondly, a hardware acceleration module for the SM4 / SM9 national cryptographic algorithms is loaded within the isolated memory space of the trusted execution environment. This module is a dedicated cryptographic processing unit, integrated within the chip or connected as a trusted peripheral. The SM4 algorithm is used for symmetric encryption, ensuring the confidentiality of sensitive information such as control commands and sensor data during transmission; the SM9 algorithm is used for asymmetric encryption and digital signatures, supporting authentication and key negotiation between devices. Due to the hardware acceleration approach, encryption and decryption operations can be completed without affecting system real-time performance, meeting the low-latency requirements of industrial control scenarios. All data involved in secure communication is automatically encrypted or decrypted upon entering or leaving the trusted execution environment, ensuring that data does not exist in plaintext form in untrusted areas.

[0080] Finally, a hardware root of trust verifies the digital signatures of critical components within the isolated memory space of the Trusted Execution Environment (TEE). During system startup or module loading, the hardware root of trust calculates the hash values ​​of the kernel, drivers, security service modules, etc., running in the TEE and compares them with pre-stored valid signatures. Only code that passes verification is allowed to be loaded and executed; any modifications that occur during storage or transmission will result in hash value mismatches, leading to system rejection. This mechanism ensures the purity and reliability of the TEE's internal operating environment, preventing malicious code from penetrating the secure area through firmware updates, remote injection, or other means.

[0081] According to one embodiment of this application, multi-source data is encrypted and collected in an isolated memory space of a trusted execution environment. The trustworthiness of the source nodes of the multi-source data is verified through remote proof, and the logical consistency between the multi-source data is analyzed to identify anomalies. The multi-source data includes network traffic characteristics, device behavior characteristics, and environmental state characteristics, specifically: Real-time capture of communication data from industrial protocols is used to encrypt and transmit network traffic characteristics to an isolated memory space within a trusted execution environment. Data from temperature, vibration, and current sensors are collected and synchronized with control commands to form a time-series data set as a characteristic of the equipment's behavior. Remote verification verifies whether the data source node is running in the isolated memory space of a trusted execution environment, preventing counterfeit devices from accessing the site. The logical relationships between multi-sensor data are analyzed within the isolated memory space of the trusted execution environment as environmental state characteristics. When the deviation exceeds a preset threshold, it is marked as suspicious data and an alarm is triggered.

[0082] As described above, firstly, communication data from protocols used in industrial control networks, such as Modbus / TCP, PROFIBUS, and OPC UA, is captured in real time. This data includes control commands between the PLC and the host computer, operating commands from the SCADA system, and equipment status feedback messages, reflecting the current interactive behavior in the network. The captured communication data, as network traffic characteristics, is encrypted at the sending end using the national cryptographic algorithm SM4 and transmitted to the Trusted Execution Environment (TEE) isolated memory space of the central security node or edge gateway. Decryption and parsing are performed within the TEE to ensure that data does not exist in plaintext in untrusted areas, preventing eavesdropping or tampering.

[0083] Secondly, physical data from various sensors on-site are collected synchronously, including key parameters such as temperature, vibration, and current. This data reflects the actual physical state of the equipment during operation. The data acquisition process is synchronized with the execution of control commands, forming a one-to-one time-series data set. For example, when a PLC issues a "start motor" command, the starting current, temperature rise rate, and vibration amplitude changes of the motor are recorded simultaneously. This time-series data set serves as a behavioral characteristic of the equipment, used to analyze the degree of matching between control actions and actual responses, and to identify whether there are any abnormalities in command execution or equipment response.

[0084] Meanwhile, to ensure the authenticity and trustworthiness of the collected data source, before data access, the recipient's trusted execution environment initiates a remote authentication request to the data sending node. The sending node generates authentication information containing a hash value of its current running state in its local trusted execution environment's isolated memory space, signs it with its private key, and returns it. The recipient verifies whether the signature and hash value match the pre-registered trusted image to determine if the recipient is running in a legitimate, tamper-free trusted environment. If verification fails, the node is considered potentially a counterfeit device or has been controlled by an attacker; the system will refuse to receive its data and issue an access alarm, effectively preventing malicious devices from impersonating legitimate nodes to access the network.

[0085] Finally, within the isolated memory space of the trusted execution environment, logical relationship analysis is performed on data from multiple sensors. Based on the physical laws and process knowledge of industrial equipment operation, a predictive correlation model between multiple sensors is established. For example, when the motor load increases, its current should rise, the vibration amplitude should increase, and the temperature should rise slowly; if situations such as "high current but no vibration" or "rapid temperature rise but no control command" occur, which violate normal logic, they are judged as abnormal. The system sets a reasonable deviation threshold (e.g., ±5%). When the actual relationship between the data from multiple sensors deviates from the expected value by more than this threshold, it is marked as suspicious data, and a security alarm is triggered, indicating that there may be sensor spoofing, data tampering, or covert attack behavior.

[0086] According to one embodiment of this application, the process of integrating network traffic characteristics, device behavior characteristics, and environmental state characteristics into a state vector, inputting it into a reinforcement learning model running in the isolated memory space of a trusted execution environment for online inference, outputting the optimal security action, and intercepting abnormal instructions in real time, specifically involves: Construct a state vector that includes network traffic characteristics, device behavior characteristics, and environmental state characteristics; A pruned and optimized deep Q-network model is run in the isolated memory space of a trusted execution environment. Forward inference is performed based on the state vector, and the expected value score of each security action is output. Select the highest-rated safety action, including logging, issuing alarms, isolating the device, or switching redundant controllers; When a PLC instruction is detected to exceed the baseline safety data by ±20%, the abnormal instruction is immediately intercepted and its transmission to the actuator is blocked.

[0087] As described above, firstly, security features from different sources are integrated into a structured state vector. This state vector serves as the input to the reinforcement learning model, comprehensively reflecting the current operational state of the industrial control system. Network traffic features, including communication frequency, message length distribution, protocol compliance, and port access behavior, characterize abnormal activity at the network layer. Device behavior features, including PLC instruction cycles, control command sequences, and operational timing patterns, reflect whether the control equipment's operation is normal. Environmental state features, including sensor data such as temperature, vibration, and current, and their changing trends, reflect the actual response of the physical process. These features are normalized within the isolated memory space of the trusted execution environment and combined in a preset order into a multi-dimensional numerical vector, forming a complete description of the current system state.

[0088] Secondly, a pruned and optimized Deep Q-Network (DQN) model is deployed and run within the isolated memory space of the Trusted Execution Environment (TEE). This model undergoes structural compression and parameter simplification, retaining key neural network connections to reduce computational resource consumption while maintaining high decision accuracy, making it suitable for industrial scenarios with limited edge resources. The model's inference process is completed within the TEE. The input is the constructed state vector, which performs forward inference operations, and the output is a set of values ​​representing the expected value score of each optional security action in the current state. These actions include "logging," "issuing an alarm," "closing communication ports," "isolating the target device," or "switching to a redundant controller," etc. A higher score indicates a greater long-term security benefit for the action in the current context.

[0089] Subsequently, based on the model's output, the system selects the security action with the highest expected value score as the optimal decision. This selection process takes place in a trusted execution environment, ensuring that the decision-making logic is not interfered with by external factors. For example, when a minor anomaly is detected but not enough to affect system operation, the model may choose to "log" and continue to observe; when obvious signs of attack are identified, it will choose to "issue an alarm" to notify operations and maintenance personnel; when a high-risk threat is determined to exist, it will directly execute strong intervention measures such as "isolate the device" or "switch to a redundant controller" to prevent the fault from spreading.

[0090] Finally, to address the most common instruction tampering attacks in industrial control systems, the system employs a threshold judgment mechanism based on a security baseline. When the system detects that the control instruction parameters received by the PLC exceed the preset security baseline data by ±20%—for example, if the set temperature is 80℃ but the instruction requires heating to 150℃, or the motor speed instruction suddenly jumps to twice the rated value—the system immediately identifies it as an abnormal instruction. This judgment result can be used as part of the state vector input model, or it can be directly triggered by the security policy module to initiate an interception action. Once a high-risk anomaly is confirmed, the system, under trusted execution environment control, immediately blocks the instruction's transmission path, preventing it from reaching the actuators (such as relays, motor drivers, etc.), thereby avoiding equipment damage or production accidents.

[0091] According to one embodiment of this application, the step of dynamically adjusting the reward function based on the feedback results of safety actions, locally updating the reinforcement learning model parameters, and aggregating the encryption gradient parameters of multiple plant areas through a federated learning mechanism to generate a global model specifically involves: Rewards are calculated based on the results of safety actions. Successful interception earns a positive reward, while false alarms or missed alarms incur negative penalties. Incremental training of the Q-network parameters of the local DQN model is performed using reward signals within the isolated memory space of the trusted execution environment, thereby improving the ability to identify attack patterns specific to the plant area. The updated gradient parameters of the local model are encrypted using an SM9 public key and a signature is attached to the isolated memory space of the trusted execution environment before being uploaded to the central server. The central server decrypts and aggregates the gradients of each node in the isolated memory space of the trusted execution environment, generates an updated global model, and encrypts and distributes it to each edge node. After verifying the global model signature, the edge nodes replace the local model parameters, thus completing the periodic evolution of the policy library.

[0092] As described above, firstly, after each security action is executed, the system collects the actual execution result of the action as feedback information and calculates a reward value accordingly. If the security action successfully blocks a confirmed attack, such as effectively intercepting illegal PLC commands or isolating infected equipment, a positive reward is given to strengthen the model's ability to identify such decisions. If a false alarm occurs, i.e., normal operation is incorrectly judged as an attack and alarms or isolation actions are executed, a negative penalty is given. If a missed alarm occurs, i.e., the actual attack is not identified or the response is insufficient, a negative penalty is also imposed. The reward value is set by classifying and quantifying factors such as the severity of the attack and the scope of business impact, forming a dynamically adjustable reward function, enabling the model to continuously adjust its learning direction based on the actual protection effect.

[0093] Secondly, within the isolated memory space of the Trusted Execution Environment (TEE) on the local edge node, the parameters of the locally deployed Deep Q-Network (DQN) model are updated using the aforementioned reward signal. This process employs incremental training, making only minor adjustments to the weight parameters of the Q-network based on the latest feedback data, thus avoiding the high computational overhead of full retraining. Since the training process is entirely completed within the TEE, model parameters, gradient information, and training data are all protected at the hardware level, preventing them from being spied on or tampered with by external programs. Through continuous local learning, the model gradually adapts to the unique equipment operation modes, process characteristics, and common attack patterns of the local plant, improving the accuracy of identifying and responding to localized threats.

[0094] Subsequently, to improve the model's generalization ability, the system uses the local learning results for global knowledge sharing. After completing the local model update, the edge nodes extract the gradient parameters generated during this training, i.e., the direction and magnitude information of the model parameter changes. This gradient data is encrypted using the public key of the SM9 algorithm, ensuring that it cannot be decrypted even if intercepted during network transmission. Simultaneously, the system attaches a digital signature to this encrypted data within the trusted execution environment to prove that the gradient originates from a legitimate and trusted device node, preventing forged nodes from uploading malicious gradients that could interfere with the global model.

[0095] The encrypted and signed gradient parameters are uploaded to the central server. The central server receives data uploaded by each edge node within its own trusted execution environment (TEE) isolated memory space. First, it verifies the digital signature of each node to confirm its legitimacy and the trustworthiness of its operating environment. Then, it decrypts the gradient parameters using an SM9 private key within the TEE. All verified gradient data are weighted and aggregated according to the data volume or trust level of each node to generate an updated global model that integrates experience from multiple plant areas. This aggregation process is completed in a trusted environment, ensuring that the global model is not contaminated or manipulated.

[0096] The generated global model, after being encrypted, is distributed to each edge node through a secure channel. Upon receiving the update package, the edge node first verifies its source signature and integrity to confirm that it is a legitimate model version from the trusted center server. After successful verification, the edge node replaces the original local model with the new model parameters within the trusted execution environment, completing the policy library update. This update process can be set to execute periodically, such as every 24 hours, or it can be dynamically triggered based on the frequency of security events.

[0097] According to one embodiment of this application, when configuration file tampering is detected, the device state is restored by calling the security baseline data stored in the isolated memory space of the trusted execution environment, and the operation log is signed and uploaded to the blockchain for evidence storage, forming a closed-loop defense ecosystem. Specifically: Continuously monitor the hash values ​​of PLC configuration programs and SCADA project files. When a discrepancy is found between the hash value and the security baseline stored in the isolated memory space of the trusted execution environment, it is determined that the file has been tampered with. The original configuration file, which is encrypted and stored in the isolated memory space of the trusted execution environment, is automatically recalled and re-burned to the PLC to restore it to a trusted operating state. The operation logs of all security events are signed with a private key in the isolated memory space of the trusted execution environment, and the signed logs are then uploaded to the blockchain.

[0098] As described above, the system continuously monitors key configuration files in the industrial control system. These files include PLC configuration programs, SCADA system engineering configuration files, and other operating parameter files that affect control logic. During monitoring, the system periodically calculates the hash value of the currently running file and compares it with the hash value of a security baseline pre-stored in the isolated memory space of the Trusted Execution Environment (TEE). This security baseline is a trusted version generated under normal equipment operation and after verification that there is no attack risk; it has immutability and a high level of trust. When the comparison result shows that the current hash value is inconsistent with the baseline, the system determines that the configuration file may have been illegally modified or maliciously tampered with, such as by implanting malicious control logic, modifying process parameters, or hiding backdoor programs.

[0099] Upon confirmation of file tampering, the system automatically initiates the recovery process. The recovery operation is triggered by the security service module within the Trusted Execution Environment (TEE), which calls an encrypted copy of the original configuration file stored in the TEE's isolated memory space. This copy was written during initial system deployment or the last security verification and is encrypted using the national cryptographic algorithm SM4, accessible only to trusted code within the TEE. The system then reprograms the decrypted original file to the target PLC or controller, overwriting the currently tampered program, restoring the equipment to a known and trusted operating state. The entire recovery process requires no manual intervention, with response time controlled within minutes, minimizing production interruptions and ensuring system availability. Simultaneously, the system generates an operation log related to this security incident. The log includes key information such as the incident time, the name of the tampered file, the original and current hash values, the type of recovery action performed, and the recovery result status. This log is generated within the isolated memory space of the Trusted Execution Environment (TEE) and immediately digitally signed using the private key bound to that device. The signing process is completed within the TEE, ensuring that the private key is not exposed to the ordinary execution environment, preventing signature forgery or key leakage. The signed log possesses non-repudiation and integrity protection; any subsequent modifications to the log content can be detected. Once signed, the operation log is uploaded to a blockchain system via a secure interface, such as an industrial security consortium blockchain built on FISCO BCOS. The upload process uses an encrypted transmission channel to ensure that data is not stolen or tampered with during transmission. The distributed ledger nature of the blockchain guarantees that once the log is uploaded, it cannot be deleted or modified, achieving long-term tamper-proof storage. This mechanism not only meets the national cybersecurity level protection system's requirement for security logs to be retained for no less than six months, but also provides a reliable basis for subsequent security audits, incident tracing, and liability determination.

[0100] Through a complete process of "monitoring—comparison—recovery—recording—evidence storage," the system not only achieves rapid remediation of attack consequences but also transforms each security incident into traceable and verifiable knowledge assets. Operations personnel can analyze attack patterns through on-chain logs, optimize security strategies, and feed typical scenarios back into the reinforcement learning model training process, further enhancing the system's ability to identify and respond to similar attacks in the future. This forms a closed-loop defense ecosystem of "perception—decision—execution—learning—optimization," enabling the continuous evolution of security capabilities.

[0101] A second aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any of the embodiments of the first aspect above.

[0102] Figure 2 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 2 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logical instructions in the memory 830 to execute the method in any of the embodiments of the first aspect described above, the method including: Activate the hardware root of trust, divide the isolated memory space of the trusted execution environment according to the hardware root of trust, and preset security baseline data in the isolated memory space of the trusted execution environment. Encrypted collection of multi-source data is performed in the isolated memory space of the trusted execution environment. The trustworthiness of the source nodes of the multi-source data is verified by remote proof. At the same time, the logical consistency between the multi-source data is analyzed to identify anomalies. The multi-source data includes network traffic characteristics, device behavior characteristics and environmental status characteristics. Network traffic characteristics, device behavior characteristics, and environmental state characteristics are integrated into a state vector, which is then input into a reinforcement learning model running in the isolated memory space of a trusted execution environment for online inference, outputting the optimal security action and intercepting abnormal commands in real time.

[0103] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0104] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer being able to perform the methods provided by the above methods, the method comprising: Activate the hardware root of trust, divide the isolated memory space of the trusted execution environment according to the hardware root of trust, and preset security baseline data in the isolated memory space of the trusted execution environment. Encrypted collection of multi-source data is performed in the isolated memory space of the trusted execution environment. The trustworthiness of the source nodes of the multi-source data is verified by remote proof. At the same time, the logical consistency between the multi-source data is analyzed to identify anomalies. The multi-source data includes network traffic characteristics, device behavior characteristics and environmental status characteristics. Network traffic characteristics, device behavior characteristics, and environmental state characteristics are integrated into a state vector, which is then input into a reinforcement learning model running in the isolated memory space of a trusted execution environment for online inference, outputting the optimal security action and intercepting abnormal commands in real time.

[0105] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the cigarette box image recognition method provided by the methods described above, the method comprising: Activate the hardware root of trust, divide the isolated memory space of the trusted execution environment according to the hardware root of trust, and preset security baseline data in the isolated memory space of the trusted execution environment. Encrypted collection of multi-source data is performed in the isolated memory space of the trusted execution environment. The trustworthiness of the source nodes of the multi-source data is verified by remote proof. At the same time, the logical consistency between the multi-source data is analyzed to identify anomalies. The multi-source data includes network traffic characteristics, device behavior characteristics and environmental status characteristics. Network traffic characteristics, device behavior characteristics, and environmental state characteristics are integrated into a state vector, which is then input into a reinforcement learning model running in the isolated memory space of a trusted execution environment for online inference, outputting the optimal security action and intercepting abnormal commands in real time.

[0106] For any parts not mentioned in this application, existing technologies may be used or referenced.

[0107] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0108] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A dynamic protection method for industrial control systems based on TEE and reinforcement learning, characterized in that, include: Activate the hardware root of trust, divide the isolated memory space of the trusted execution environment according to the hardware root of trust, and preset security baseline data in the isolated memory space of the trusted execution environment. Encrypted collection of multi-source data is performed in the isolated memory space of the trusted execution environment. The trustworthiness of the source nodes of the multi-source data is verified by remote proof. At the same time, the logical consistency between the multi-source data is analyzed to identify anomalies. The multi-source data includes network traffic characteristics, device behavior characteristics and environmental status characteristics. Network traffic characteristics, device behavior characteristics, and environmental state characteristics are integrated into a state vector, which is then input into a reinforcement learning model running in the isolated memory space of a trusted execution environment for online inference, outputting the optimal security action and intercepting abnormal commands in real time.

2. The method according to claim 1, characterized in that, Also includes: The reward function is dynamically adjusted based on the feedback results of safety actions, the parameters of the reinforcement learning model are locally updated, and the encrypted gradient parameters of multiple plant areas are aggregated through a federated learning mechanism to generate a global model.

3. The method according to claim 1, characterized in that, Also includes: When configuration file tampering is detected, the device state is restored by calling the security baseline data stored in the isolated memory space of the trusted execution environment, and the operation log is signed and stored on the blockchain to form a closed-loop defense ecosystem.

4. The method according to claim 1, characterized in that, The activation of the hardware trust root involves dividing the trusted execution environment into isolated memory spaces based on the hardware trust root, and pre-setting security baseline data within these isolated memory spaces. Specifically: Activate the hardware root of trust and generate a unique hardware identifier for the device through a physically unclonable function for identity authentication and integrity verification; A hardware acceleration module for the SM4 / SM9 national cryptographic algorithm is loaded within the isolated memory space of the trusted execution environment to achieve real-time encrypted transmission of control commands and sensor data. The digital signature within the isolated memory space of the trusted execution environment is verified through a hardware root of trust to ensure it has not been tampered with.

5. The method according to claim 1, characterized in that, Encrypted data from multiple sources is collected in an isolated memory space within a trusted execution environment. The trustworthiness of the source nodes of the multi-source data is verified through remote authentication. Simultaneously, the logical consistency between the multi-source data is analyzed to identify anomalies. The multi-source data includes network traffic characteristics, device behavior characteristics, and environmental state characteristics, specifically: Real-time capture of communication data from industrial protocols is used to encrypt and transmit network traffic characteristics to an isolated memory space within a trusted execution environment. Data from temperature, vibration, and current sensors are collected and synchronized with control commands to form a time-series data set as a characteristic of the equipment's behavior. Remote verification verifies whether the data source node is running in the isolated memory space of a trusted execution environment, preventing counterfeit devices from accessing the site. The logical relationships between multi-sensor data are analyzed within the isolated memory space of the trusted execution environment as environmental state characteristics. When the deviation exceeds a preset threshold, it is marked as suspicious data and an alarm is triggered.

6. The method according to claim 1, characterized in that, The process involves integrating network traffic characteristics, device behavior characteristics, and environmental state characteristics into a state vector, inputting it into a reinforcement learning model running in the isolated memory space of a trusted execution environment for online inference, outputting the optimal security action, and intercepting abnormal commands in real time. Specifically: Construct a state vector that includes network traffic characteristics, device behavior characteristics, and environmental state characteristics; A pruned and optimized deep Q-network model is run in the isolated memory space of a trusted execution environment. Forward inference is performed based on the state vector, and the expected value score of each security action is output. Select the highest-rated safety action, including logging, issuing alarms, isolating the device, or switching redundant controllers; When a PLC instruction is detected to exceed the baseline safety data by ±20%, the abnormal instruction is immediately intercepted and its transmission to the actuator is blocked.

7. The method according to claim 2, characterized in that, The process involves dynamically adjusting the reward function based on feedback from safety actions, locally updating the reinforcement learning model parameters, and aggregating encryption gradient parameters from multiple plant areas through a federated learning mechanism to generate a global model. Specifically: Rewards are calculated based on the results of safety actions. Successful interception earns a positive reward, while false alarms or missed alarms incur negative penalties. Incremental training of the Q-network parameters of the local DQN model is performed using reward signals within the isolated memory space of the trusted execution environment, thereby improving the ability to identify attack patterns specific to the plant area. The updated gradient parameters of the local model are encrypted using an SM9 public key and a signature is attached to the isolated memory space of the trusted execution environment before being uploaded to the central server. The central server decrypts and aggregates the gradients of each node in the isolated memory space of the trusted execution environment, generates an updated global model, and encrypts and distributes it to each edge node. After verifying the global model signature, the edge nodes replace the local model parameters, thus completing the periodic evolution of the policy library.

8. The method according to claim 3, characterized in that, When configuration file tampering is detected, the device state is restored by calling the security baseline data stored in the isolated memory space of the trusted execution environment, and the operation log is signed and uploaded to the blockchain for evidence storage, forming a closed-loop defense ecosystem. Specifically: Continuously monitor the hash values ​​of PLC configuration programs and SCADA project files. When a discrepancy is found between the hash value and the security baseline stored in the isolated memory space of the trusted execution environment, it is determined that the file has been tampered with. The original configuration file, which is encrypted and stored in the isolated memory space of the trusted execution environment, is automatically recalled and re-burned to the PLC to restore it to a trusted operating state. The operation logs of all security events are signed with a private key in the isolated memory space of the trusted execution environment, and the signed logs are then uploaded to the blockchain.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1-8.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-8.