An intelligent control system for devices based on the Internet of Things
Patent Information
- Application Number
- CN202611072471.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]发明人经研究发现,上述两类现有技术方案均存在难以克服的固有缺陷
[0015]1.本发明通过在边缘侧引入基于多智能体深度强化学习的确定性时空联合调度机制,实时感知控制流、感知流、交互流的动态变化,动态划分网络时隙资源与边缘节点计算算力,从而消除产线柔性化重构、交互流突发场景下对关键控制流造成的确定性延迟破坏,相较于依赖离线静态调度表的现有方案,能够在动态工况下持续保障关键反馈闭环控制流的低延迟传输。
Smart Images

Figure CN122824769A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial Internet of Things (IoT) technology, specifically to an IoT-based intelligent control system for equipment, and more particularly to an intelligent control system for industrial field equipment that integrates edge computing, zero-trust security protection, and deterministic network scheduling. It can be applied to cloud-edge collaborative control of factory production lines, secure access to legacy controllers, and deterministic transmission assurance for heterogeneous business flows. Background Technology
[0002] With the in-depth development of industrial IoT and smart manufacturing technologies, cloud-edge collaborative intelligent equipment control systems have become one of the key technical paths to achieve production line flexibility and intelligence. In such systems, data generated by field equipment is usually aggregated and processed by edge nodes before deciding whether to upload it to the cloud. At the same time, cloud control commands also need to be sent to field execution agencies through edge nodes, forming an overall closed loop consisting of cloud decision-making, edge execution and field feedback.
[0003] Modern industrial field networks simultaneously contain multiple types of traffic flows with distinct characteristics: control traffic flows, characterized by high priority, low transmission latency, and deterministic cycles, directly impacting production safety and product quality; sensing traffic flows, characterized by periodic data collection and large data volumes, used for equipment status monitoring and process parameter acquisition; and interactive traffic flows, characterized by sporadic and sudden occurrences, typically generated by cloud-based control commands or human-machine interaction. To address the mixed transmission problem of these heterogeneous traffic flows, the industry has widely adopted Time-Sensitive Networking (TSN) standards. Through time synchronization, traffic shaping, and static scheduling table mechanisms, deterministic transmission guarantees are provided for traffic flows of different priorities. The implementation steps of this type of solution are typically as follows: first, the network management entity synchronizes the time of all network nodes; second, based on the pre-planned traffic flow priorities and cycles, a static gating scheduling table is generated during the network configuration phase and distributed to each switching node; during network operation, each node releases packets of corresponding priorities within a specific time window according to the pre-configured scheduling table, thereby achieving deterministic transmission.
[0004] On the other hand, factories still commonly have older, existing controllers with long design lifespans. These devices often use fieldbus protocols with open message structures and low protocol overhead, but generally lack authentication and encryption mechanisms for data interaction. To address the lack of security protection for such protocols, existing solutions typically use a protocol conversion gateway combined with traffic statistics and intrusion detection. Its structure usually includes: a protocol conversion module, which converts old protocol messages into standard Ethernet protocols or cloud-recognizable formats to enable interconnection between heterogeneous devices; and a traffic statistics and analysis module, which establishes a normal traffic baseline by collecting statistical characteristics of network traffic and compares real-time traffic with this baseline. If the traffic exceeds a preset threshold, it is judged as abnormal and an alarm is triggered. The above two modules are usually deployed in series in industrial gateway devices.
[0005] The inventors discovered through research that both of the above-mentioned existing technical solutions have inherent defects that are difficult to overcome.
[0006] Regarding deterministic scheduling schemes based on time-sensitive networking standards, since the scheduling table is generated offline during the network configuration phase based on pre-planned fixed service flow characteristics, when scenarios such as plug-and-play equipment and flexible reconfiguration occur on the production line, causing dynamic changes in service flow characteristics, the original scheduling table cannot detect these changes and update them in a timely manner. Because the scheduling table cannot be updated in a timely manner, when there is a sudden surge in interactive service flows, this sudden traffic will occupy the time slot resources originally allocated to the control service flows. As the time slot resources of the control service flows are squeezed, the queuing delay of the critical feedback closed-loop control flow will inevitably deteriorate, ultimately causing the deterministic control guarantee to fail under dynamic operating conditions.
[0007] Regarding the security protection scheme based on protocol conversion gateway plus traffic statistics intrusion detection, since the protocol conversion module only converts the message format without changing the transmission method and content verification mechanism of the message itself, the security shortcomings of plaintext transmission of old protocols and lack of identity authentication are not substantially compensated at this stage. Since security protection relies entirely on the traffic statistics analysis module, which only judges whether the traffic is abnormal based on macroscopic statistical characteristics, without deeply analyzing the specific control semantics carried by the message and whether it conforms to the physical laws of the process, this scheme is difficult to identify attack messages with completely legal forged formats that only tamper with specific control parameters because their statistical characteristics are highly similar to normal traffic, and the protection capability fails. At the same time, since normal operating conditions in industrial sites have reasonable changes such as large load fluctuations and batch switching, these reasonable changes are likely to exceed the preset threshold in terms of statistical characteristics, causing this scheme to generate a large number of false alarms, thus limiting its practical value.
[0008] In summary, how to achieve dynamic adaptive scheduling of network time slot resources and edge computing power in a heterogeneous network environment where control flow, sensing flow and interaction flow are mixed in industrial field, and how to provide effective and low false alarm zero-trust security protection for existing legacy protocol controllers without modification, shutdown, or firmware flashing are the technical problems that urgently need to be solved in this field. Summary of the Invention
[0009] The technical problems to be solved by this invention include the following two aspects: First, how to achieve dynamic and adaptive scheduling of network time slot resources and edge computing power in a heterogeneous network environment where control business flow, sensing business flow and interactive business flow are mixed and transmitted in industrial field, so as to ensure the deterministic and low-latency transmission of key feedback closed-loop control flow under complex working conditions such as flexible reconstruction of production lines and sudden bursts of interactive traffic, and avoid the failure of existing deterministic network fixed configuration schemes in dynamic scenarios. Second, how to provide effective zero-trust security protection for existing controllers that use a large number of unauthenticated plaintext transmission protocols in factories without modification, shutdown, or firmware flashing, fundamentally identify and block control commands with forged legitimate formats, and at the same time avoid the problem of a large number of false alarms generated by existing intrusion detection schemes based on traffic statistics baselines due to normal process fluctuations.
[0010] To address the aforementioned technical problems, this invention provides an IoT-based intelligent device control system, comprising a cloud-based training and management layer, an edge intelligent control layer, and a field device layer. The field device layer includes intelligent devices supporting standard protocols and existing controllers using outdated industrial protocols. The field device layer and the edge intelligent control layer are connected via an industrial fieldbus or an industrial Ethernet. The edge intelligent control layer is deployed on an edge computing node and includes a protocol-aware zero-trust security proxy module, a heterogeneous business flow deterministic spatiotemporal joint scheduling module, and a resource arbitration unit.
[0011] The protocol-aware zero-trust security proxy module includes a message parsing unit, a knowledge graph inference unit, and an anomaly determination unit connected in sequence. The message parsing unit is used to parse industrial protocol messages flowing through the network interface card (NIC) of the edge computing node without changing the old industrial protocol message format, to obtain instruction semantic information including protocol type, function code, target register address, and data value to be written. The knowledge graph inference unit stores a process knowledge graph, which includes a set of nodes representing physical quantities and a set of constraint edges representing the mechanistic constraint relationships between physical quantities. The knowledge graph inference unit is used to map the target register address in the instruction semantic information to the corresponding physical quantity node according to the preset mapping relationship between register addresses and graph nodes, and substitute the data value to be written into the physical constraint equation defined by the constraint edge connected to the physical quantity node to obtain the deviation of the instruction relative to each constraint edge. The anomaly determination unit is used to calculate a comprehensive anomaly score based on the deviation of each constraint edge and a preset weight, and compare the comprehensive anomaly score with a preset determination threshold, thereby deciding whether to forward the instruction to the target stock controller or block it.
[0012] The heterogeneous service flow deterministic spatiotemporal joint scheduling module includes a service flow feature perception unit, a policy reasoning unit, and a scheduling configuration unit connected in sequence. The service flow feature perception unit is used to collect the queue length, arrival rate, and remaining deadline of the control service flow, the perception service flow, and the interactive service flow in the current scheduling period to obtain a state observation vector. The policy reasoning unit includes a time slot allocation agent and a computing power allocation agent trained based on a multi-agent deep reinforcement learning architecture. The two share the state observation vector as input and output the allocation ratio of network transmission time slots among the three types of service flows, and the allocation ratio of edge node computing resources among protocol parsing, knowledge graph reasoning, and scheduling decision tasks, respectively. The scheduling configuration unit is used to generate flow control configuration parameters and computing resource quota parameters according to the above allocation ratios and distribute them to the network interface and operating system resource scheduling interface of the edge nodes for execution.
[0013] The resource arbitration unit is used to coordinate the use of the protocol-aware zero-trust security proxy module and the heterogeneous business flow deterministic spatiotemporal joint scheduling module for edge node network interfaces and computing resource pools, so that the two are connected and coordinated on the same data path, rather than simply stacking independent functional modules.
[0014] The present invention can achieve the following beneficial effects:
[0015] 1. This invention introduces a deterministic spatiotemporal joint scheduling mechanism based on multi-agent deep reinforcement learning at the edge side to perceive the dynamic changes of control flow, perception flow, and interaction flow in real time, and dynamically allocate network time slot resources and edge node computing power. This eliminates the deterministic delay damage to critical control flow caused by flexible reconfiguration of production lines and sudden interaction flow scenarios. Compared with existing solutions that rely on offline static scheduling tables, this invention can continuously ensure low-latency transmission of critical feedback closed-loop control flow under dynamic operating conditions.
[0016] 2. This invention deploys a protocol-agnostic zero-trust security proxy mechanism in series at the edge, utilizes an extended Berkeley packet filtering program to perform line-rate deep parsing of legacy protocol packets at the network interface card level, and combines a process knowledge graph containing physical mechanism constraint edges to perform semantic-level anomaly detection of issued commands. This achieves zero-trust security protection for legacy protocols without requiring any downtime modifications or firmware upgrades to existing controllers. Because the judgment is based on process physical laws rather than traffic statistics, it can identify attack commands with legitimate formats but altered control parameters, while avoiding misjudging reasonable traffic changes caused by normal process fluctuations as anomalies. Compared to intrusion detection schemes that rely on statistical baselines, this invention performs semantic-level judgment based on process physical mechanism constraints, enabling it to identify attack commands with legitimate formats but altered control parameters, and continuously reduces the false alarm rate through graph state updates and feedback iteration mechanisms.
[0017] 3. This invention deploys the security proxy mechanism and the deterministic scheduling mechanism on the same edge computing node, sharing network interfaces and computing resource pools. It also coordinates the allocation of resources through a resource arbitration unit, enabling zero-trust security protection and deterministic transmission assurance to be seamlessly integrated and coordinated on the same data path. This avoids the performance constraints caused by the independent deployment and resource contention of the two types of functional modules, thereby improving the overall resource utilization efficiency and engineering feasibility of the system. Attached Figure Description
[0018] Figure 1 This is a flowchart of the protocol's seamless zero-trust security proxy module processing flow.
[0019] Figure 2 This is a flowchart of the heterogeneous service flow deterministic spatiotemporal joint scheduling module of the present invention;
[0020] Figure 3 This is a flowchart illustrating the overall collaborative workflow of the system of this invention. Detailed Implementation
[0021] The following is in conjunction with the appendix Figure 1-3The specific embodiments of the present invention will be further described below. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0022] I. System Overall Architecture
[0023] The IoT-based intelligent control system provided in this embodiment adopts a three-layer architecture of cloud-edge collaboration, consisting of a cloud training and management layer, an edge intelligent control layer, and a field device layer from top to bottom.
[0024] The cloud-based training and management layer is deployed on a cloud server cluster. It contains a historical business flow database, an offline training module for multi-agent deep reinforcement learning models, and a module for constructing and updating process knowledge graphs. This layer distributes the trained lightweight scheduling and inference model and the updated process knowledge graph to the edge intelligent control layer through a cloud-edge communication link. At the same time, it receives de-identified running data uploaded by the edge layer for continuous iterative optimization of the model.
[0025] The edge intelligent control layer is deployed on edge computing nodes in industrial sites. These edge computing nodes can be edge gateways or edge controllers implemented based on processor architectures. The layer has two main functional modules arranged in parallel: one is a protocol-aware zero-trust security proxy module, which consists of three sub-units connected in series: a message parsing unit, a knowledge graph reasoning unit, and an anomaly judgment unit; the other is a heterogeneous business flow deterministic spatiotemporal joint scheduling module, which consists of three sub-units connected in series: a business flow feature perception unit, a policy reasoning unit, and a scheduling configuration unit. The above two functional modules share the network interface and computing resource pool of the edge computing node and coordinate their resource usage through a resource arbitration unit.
[0026] The field device layer includes newly added smart devices that support standard protocols, as well as existing controllers, sensors, and actuators that use outdated industrial protocols. The field device layer and the edge intelligent control layer are connected via industrial fieldbus or industrial Ethernet. Control commands are sent down from the cloud or edge to the field devices, and operating data is uploaded from the field devices to the edge or cloud.
[0027] II. Protocol-Invisible Zero-Trust Security Proxy Mechanism
[0028] The core of this mechanism is to not change the message format and transmission method of the old protocol itself, but to insert a security proxy in a series manner during the data flow through the network card of the edge computing node to perform deep parsing and semantic verification of the message.
[0029] In the process knowledge graph construction phase, for the specific process of the target production line, the mechanistic constraint relationships between key physical quantities are sorted out. Each physical quantity is abstracted as a node in the graph, and the mechanistic constraint equations between physical quantities are abstracted as constraint edges connecting the corresponding nodes, forming a process knowledge graph containing a set of nodes and a set of constraint edges. This graph is stored locally on the edge computing nodes. Physical quantity nodes can include process parameters such as temperature, pressure, flow rate and valve opening, and constraint edges can include edges that characterize the mechanistic relationships between physical quantities, such as temperature rise rate constraint edges and pressure-flow rate correlation constraint edges.
[0030] During the line-rate parsing phase, an extended Berkeley packet filter program, pre-written and loaded into the network card driver's transmit / receive path, parses each passing industrial protocol message byte by byte, extracting the message's protocol type, function code, destination register address, and the data value to be written, forming structured instruction semantic information, which can be represented as a quintuple. ,in This indicates the timestamp of the instruction. Indicates the protocol type. Indicates function code, Indicates the address of the target register. This represents the data value to be written.
[0031] In the semantic mapping stage, based on the pre-established mapping table between register addresses and process knowledge graph nodes, the extracted target register addresses are... Map the data to the corresponding physical quantity node in the process knowledge graph to determine the specific physical quantity that the instruction intends to modify and its current historical state value recorded in the graph.
[0032] During the physical constraint verification phase, starting from the mapped physical quantity node, all constraint edges connected to that node in the process knowledge graph are traversed, and the instruction data values are... Substitute the physical constraint equations defined by each constraint edge into the equations, calculate the degree of satisfaction of each constraint equation, and obtain the result of the instruction relative to the first constraint edge. Deviation of the constrained edge Taking the temperature rise rate constraint of temperature-related physical quantities as an example, its constraint equation can be expressed as:
[0033] ;
[0034] in, This indicates the target temperature value that the current command wants to write. This represents the current state value of the physical quantity of temperature recorded in the process knowledge graph. This represents the time interval from the last update of the state of this physical quantity to the current time when the command is issued. This indicates the upper limit of the allowable heating rate of the physical quantity of temperature within the safe operating range of the equipment. It is predetermined by the process mechanism. If this inequality does not hold, it means that the temperature change rate required by the current instruction exceeds the physically achievable or safe range, and the deviation should be taken into account. .
[0035] Taking the physical relationship between pressure and flow velocity as another example, it can be characterized using a simplified Bernoulli relation:
[0036] ;
[0037] in, This indicates the target pressure value that the current instruction intends to write. This represents the density of the medium, which is a known process parameter. This represents the current flow rate state value associated with this pressure node, recorded in the process knowledge graph. This represents the Bernoulli constant of the pipeline system in steady state, which is pre-calibrated using historical steady-state operating data. This indicates the allowable engineering error margin. If this inequality does not hold, it means that the pressure value required by the instruction is physically inconsistent with the current flow rate, and the deviation should be taken into account. .
[0038] In the comprehensive anomaly score calculation and judgment stage, the deviations of each constraint edge are weighted and summed according to preset weights to obtain the comprehensive anomaly score of the instruction:
[0039] ;
[0040] in, This represents the total number of connected constraint edges for the physical quantity node in the process knowledge graph. Indicates the first The preset weights corresponding to the constraint edges are used to reflect the differences in the importance of different constraints in terms of security. This indicates that the instruction obtained from the aforementioned calculation is relative to the first... The deviation of each constraint edge will be used to calculate the overall anomaly score. Compared with the preset judgment threshold Comparison: If Less than If the instruction is deemed normal, it is allowed and forwarded to the target inventory controller. Not less than If the instruction is deemed abnormal, it will be blocked, and the complete semantic information of the abnormal instruction will be reported to the cloud management layer, while triggering a local alarm.
[0041] During the graph status update and feedback iteration phase, for instructions that are judged to be normal and released, the target physical quantity value carried by them is updated to the historical status of the corresponding node in the process knowledge graph. For alarm records that are confirmed as misjudged in the cloud, the weight of the corresponding constraint edge or the judgment threshold parameter is fed back to the aforementioned judgment weight configuration for dynamic correction, so as to continuously reduce the false alarm rate.
[0042] III. Deterministic Spatiotemporal Joint Scheduling Mechanism for Heterogeneous Service Flows
[0043] This mechanism uses the edge side to perceive the dynamic characteristics of three types of service flows—control flow, perception flow, and interaction flow—in real time, and uses a multi-agent deep reinforcement learning algorithm to dynamically output the allocation strategy of network time slot resources and computing power, replacing the traditional static scheduling table.
[0044] During the state space construction phase, at the beginning of each scheduling cycle, the business flow feature perception unit collects information such as the current queue length, arrival rate, and remaining deadline of the control flow, perception flow, and interaction flow to construct the system's state vector for that cycle. It serves as the common observation input for both the time slot allocation agent and the computing power allocation agent.
[0045] In the action space definition phase, the action of the time slot allocation agent is defined as the allocation ratio vector of network transmission time slots among the three types of service flows in the next scheduling cycle. The action of the computing power allocation agent is defined as the allocation ratio vector of edge node computing resources among tasks such as protocol parsing, knowledge graph reasoning, and scheduling decision-making. Both agents, based on the current policy network and state vectors, [follow these steps]. Output their respective actions and .
[0046] During the joint action execution phase, the two actions mentioned above are merged into a joint scheduling configuration. The scheduling configuration unit converts the time slot allocation ratio into specific queuing rules and queue weight parameters and sends them to the network interface. At the same time, the computing power allocation ratio is converted into specific central processing unit resource quotas and sent to the resource scheduling interface of the edge node operating system, thus completing the actual resource reconstruction for this cycle.
[0047] During the reward signal calculation phase, after the joint action is completed, the actual end-to-end delay of the control flow, the packet loss rate of the sensing flow, and the average waiting time of the interaction flow within this cycle are collected. The immediate reward for this cycle is then calculated according to the preset reward function.
[0048] ;
[0049] in, This indicates the extent to which the actual end-to-end latency of the control service flow exceeds its deterministic latency threshold within the current cycle. This indicates the packet loss rate of the sensing business flow during this period. This indicates the average waiting time of the interactive service flow within this period. , , These represent the weight coefficients of the three indicators in the reward function, and Value greater than and The value of is determined to reflect the priority of control flow determinism in resource allocation.
[0050] During the policy network update phase, the state of this period will be updated. ,action , ,award and the state of the next cycle The constructed experience tuples are stored in the experience replay pool. The cloud training module periodically extracts experience tuples in batches from the experience replay pool. A multi-agent actor-commentator framework is used to update the gradients of the policy network parameters and value network parameters of the two agents respectively. The policy network is updated using the policy gradient method under the actor-commentator framework. The temporal difference error of the value network is defined as follows:
[0051] ;
[0052] in, The immediate reward obtained from the aforementioned calculation, This represents the discount factor, used to weigh the relative importance of current rewards and future rewards. Representing the value network in relation to state The value estimate, This indicates the value network's state for the next cycle. The value estimate, This is the temporal difference error, used to guide the gradient update direction of the aforementioned policy network and value network parameters. The cloud training module distributes the updated lightweight policy network parameters to the edge computing nodes to replace the original inference model and complete one iteration.
[0053] During the edge-side loop inference execution phase, when the edge computing node does not receive updates from the cloud, it continuously utilizes the currently issued policy network parameters to repeatedly execute the aforementioned state awareness, action output, and joint execution process for each new scheduling cycle. This enables local closed-loop scheduling decisions without relying on real-time cloud responses, thereby ensuring low latency in scheduling decisions and continuous availability of the system in the event of cloud-edge communication interruptions.
[0054] IV. Collaborative Work Sequence
[0055] Once a control command is generated from the field controller or the cloud, it first enters the protocol-aware, zero-trust security proxy module of the edge computing node. Here, message parsing, knowledge graph comparison, and anomaly detection are completed sequentially. For commands that are determined to be normal, they further enter the heterogeneous business flow deterministic spatiotemporal joint scheduling module. Here, state awareness, policy reasoning, and time slot and computing power configuration are completed sequentially. Finally, the command is actually issued and executed. The above message parsing, semantic judgment, and deterministic scheduling links are closely linked in time and are on the same data path. This demonstrates that the security protection and deterministic transmission guarantee of this system are a holistic technical solution that works together, rather than a simple stacking of two independent functional modules.
[0056] V. Other Implementation Methods
[0057] In another embodiment of the present invention, the knowledge graph reasoning unit in the protocol-aware zero-trust security proxy module can also be replaced by a physical constraint verification unit based on a rule engine. That is, the mechanism constraint relationship of each physical quantity is pre-written into a series of explicit conditions and action rules by process experts, and a lightweight rule engine is deployed on the edge computing node to perform legality verification on each rule of the instruction semantic information parsed by the message parsing unit. This can also achieve semantic-level anomaly determination without relying on traffic statistics features.
[0058] In another embodiment of the present invention, the strategy reasoning unit in the heterogeneous service flow deterministic spatiotemporal joint scheduling module can also be replaced by a scheme based on a service flow prediction model of a long short-term memory network combined with an optimization algorithm. That is, the long short-term memory network model predicts the arrival characteristics of the three types of service flows in the next scheduling cycle based on historical flow data, and then inputs the prediction results into a preset resource allocation optimization objective function to solve for the allocation scheme of time slots and computing power, which can also realize the dynamic adjustment of resources.
[0059] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, any modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the protection scope defined by the claims of the present invention.
Claims
1. An IoT-based intelligent device control system, comprising a cloud-based training and management layer, an edge intelligent control layer, and a field device layer, wherein the field device layer includes intelligent devices supporting standard protocols and existing controllers using legacy industrial protocols, and the field device layer and the edge intelligent control layer are connected via an industrial fieldbus or an industrial Ethernet, characterized in that, The edge intelligent control layer is deployed on edge computing nodes and includes: a protocol-aware zero-trust security proxy module, a heterogeneous business flow deterministic spatiotemporal joint scheduling module, and a resource arbitration unit; The protocol-aware zero-trust security proxy module includes a message parsing unit, a knowledge graph reasoning unit, and an anomaly determination unit connected in sequence. The message parsing unit is used to parse the industrial protocol messages that flow through the network card of the edge computing node without changing the old industrial protocol message format, and obtain instruction semantic information including protocol type, function code, target register address and data value to be written. The knowledge graph reasoning unit stores a process knowledge graph, which includes a set of nodes representing physical quantities and a set of constraint edges representing the mechanistic constraint relationship between physical quantities. The knowledge graph reasoning unit is used to map the target register address in the instruction semantic information to the corresponding physical quantity node according to the preset mapping relationship between register address and graph node, and substitute the data value to be written into the physical constraint equation defined by the constraint edge connected to the physical quantity node to obtain the deviation of the instruction relative to each constraint edge. The anomaly determination unit is used to calculate a comprehensive anomaly score based on the deviation of each constraint edge and a preset weight, compare the comprehensive anomaly score with a preset determination threshold, and determine that the instruction is a normal instruction and forward it to the target stock controller when the comprehensive anomaly score is less than the preset determination threshold. When the comprehensive anomaly score is not less than the preset determination threshold, the instruction is determined to be an anomaly instruction and blocked. The heterogeneous service flow deterministic spatiotemporal joint scheduling module includes a service flow feature perception unit, a strategy reasoning unit, and a scheduling configuration unit connected in sequence. The service flow feature perception unit is used to collect the queue length, arrival rate, and remaining deadline of the control service flow, perception service flow, and interactive service flow in the current scheduling cycle to obtain a state observation vector. The policy reasoning unit includes a time slot allocation agent and a computing power allocation agent trained based on a multi-agent deep reinforcement learning architecture. The time slot allocation agent and the computing power allocation agent share the state observation vector as input and output the allocation ratio of network transmission time slots among the control service flow, the perception service flow and the interaction service flow, respectively, as well as the allocation ratio of edge node computing resources among protocol parsing, knowledge graph reasoning and scheduling decision tasks. The scheduling configuration unit is used to generate flow control configuration parameters and computing resource quota parameters according to the allocation ratio, and distribute them to the network interface and operating system resource scheduling interface of the edge node, respectively. The resource arbitration unit is used to prioritize exclusive access to the network interface send / receive buffer and processor core resources during the processing period when the protocol-aware zero-trust security proxy module parses and judges the message, according to the preset resource occupation priority rules. After the anomaly judgment is completed, the resource occupation permission is released and handed over to the heterogeneous service flow deterministic spatiotemporal joint scheduling module for time slot and computing power configuration, thereby coordinating the occupation of the edge node network interface and computing resource pool by the two modules.
2. The system according to claim 1, characterized in that, The message parsing unit includes an extended Berkeley packet filtering program loaded into the network card driver's receive path and send path. The extended Berkeley packet filtering program is used to parse each industrial protocol message that passes through the network card byte by byte.
3. The system according to claim 1, characterized in that, The anomaly determination unit is also used to update the data value to be written carried by the instruction that is determined to be normal to the historical state of the corresponding physical quantity node in the process knowledge graph. For the abnormal instruction that is confirmed as a misjudgment by the cloud management layer, the preset weight of the corresponding constraint edge and at least one of the preset determination threshold are fed back for correction.
4. The system according to claim 1, characterized in that, The physical constraint equations defined by the constraint edges are the heating rate constraint equation, which characterizes the rate of change of physical quantities, and the pressure-flow rate correlation constraint equation, which characterizes the relationship between multiple physical quantities.
5. The system according to claim 1, characterized in that, The cloud-based training and management layer includes a historical business flow database, an offline training module for multi-agent deep reinforcement learning models, and a process knowledge graph construction and update module. It is used to iteratively update the policy network parameters of the time slot allocation agent and the computing power allocation agent based on the de-identified operating data uploaded by the edge intelligent control layer, and to send the updated policy network parameters to the policy inference unit.
6. The system according to claim 1, characterized in that, The scheduling configuration unit is used to convert the allocation ratio into queue weight parameters of flow control queuing rules and send them to the network interface, and to convert them into central processing unit resource quota parameters and send them to the operating system resource scheduling interface.
7. The system according to claim 1, characterized in that, It also includes a reward calculation unit, which is used to collect the actual end-to-end latency of the control service flow, the packet loss rate of the sensing service flow, and the average waiting time of the interactive service flow within the current scheduling period, and calculate the instant reward according to a preset reward function. In the reward function, the weight coefficient corresponding to the control service flow is greater than the weight coefficient corresponding to the sensing service flow, and is also greater than the weight coefficient corresponding to the interactive service flow.
8. The system according to claim 1, characterized in that, The edge intelligent control layer is also used to continuously use the currently issued policy network parameters to make local closed-loop scheduling decisions for each new scheduling cycle when no policy network parameter update is received from the cloud training and management layer.
9. A device intelligent control method based on the Internet of Things, applied to a system including a cloud training and management layer, an edge intelligent control layer, and a field device layer, characterized in that, include: The network interface card (NIC) at the edge intelligent control layer parses the industrial protocol messages that flow through it, and obtains instruction semantic information including protocol type, function code, target register address and data value to be written. Based on the preset mapping relationship between register addresses and graph nodes, the target register address is mapped to the corresponding physical quantity node in the process knowledge graph, and the data value to be written is substituted into the physical constraint equation defined by the constraint edge connected to the physical quantity node to obtain the deviation of the instruction relative to each constraint edge. A comprehensive anomaly score is calculated based on the deviation of each constraint edge and a preset weight. The comprehensive anomaly score is compared with a preset judgment threshold. When the comprehensive anomaly score is less than the preset judgment threshold, the instruction is forwarded to the target stock controller. When the comprehensive anomaly score is not less than the preset judgment threshold, the instruction is blocked. Collect the queue length, arrival rate and remaining deadline of the control service flow, sensing service flow and interactive service flow in the current scheduling cycle to obtain the state observation vector; The state observation vector is input into the time slot allocation agent and the computing power allocation agent respectively to obtain the allocation ratio of network transmission time slots among the control service flow, the perception service flow and the interaction service flow, as well as the allocation ratio of edge node computing resources among protocol parsing, knowledge graph reasoning and scheduling decision tasks. Based on the allocation ratio, flow control configuration parameters and computing resource quota parameters are generated and respectively sent to the network interface and operating system resource scheduling interface of the edge node for execution. The resource arbitration unit coordinates the parsing and judgment process of the instruction and the execution process of the allocation ratio to address the shared use of the edge node network interface and computing resource pool.
10. The method according to claim 9, characterized in that, Also includes: The system collects the actual end-to-end latency of the control service flow, the packet loss rate of the sensing service flow, and the average waiting time of the interactive service flow within the current scheduling cycle. It then calculates an immediate reward and updates the policy network parameters of the time slot allocation agent and the computing power allocation agent based on the immediate reward.