Method and system for dynamic control of internet of things devices based on cloud collaboration

CN122601673APending Publication Date: 2026-08-18RECREAT AIOT (GD) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610708201.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]在目前的云端控制系统中,将物联网设备上报的传感数据与网络状态数据进行孤立处理或简单的特征拼接,未能构建统一的全局状态空间以提取包含时序依赖与空间耦合的动态特征向量,这种扁平化的特征处理方式忽略了设备集群之间物理拓扑与逻辑通信的深层关联,使得云端对系统态势的把控存在盲区,同时,现有的控制偏移度计算多基于当前时刻的瞬时状态阈值进行判定,缺乏在云端协同空间中对未来状态轨迹的主动预演,影响了控制偏移度的精准性,导致了物联网设备的动态控制内容的精准性较低

Benefits of technology

(1)在云服务器中,接收物联网设备的多源异构流数据,并结合预设的时空关联图构建物联网设备的全局状态空间,对全局状态空间进行特征聚合与降维,提取出包含时序依赖特性与空间耦合特性的动态特征向量;将动态特征向量输入对应的云端协同空间,并在特征演变过程中生成物联网设备在未来时域内的状态预演轨迹,并结合预设的安全运行轨迹的对比而确定对应的控制偏移度,引入了动态特征向量,对云端协同空间进一步把控,提高了控制偏移度的精准性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122601673A_ABST
    Figure CN122601673A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on cloud end cooperation's dynamic control method and system of internet of things equipment, and the application relates to the technical field of cloud end cooperation, in cloud server, dynamic feature vector is input corresponding cloud end cooperation space, and the state pre-play trajectory of internet of things equipment in future time domain is generated in feature evolution process, and the control offset degree of correspondence is determined by combining the comparison of preset safe operation trajectory, improve the accuracy of control offset degree.Cloud server determines the cooperative control level matched with optimal control strategy based on dynamic feature vector and state evolution trajectory;Cloud server issues optimal control strategy to corresponding cooperative control node according to cooperative control level, and real-time acquisition observation data after strategy execution, further determine the dynamic control content of internet of things equipment in combination with the current load condition of cloud server and the current state of internet of things equipment, improve the accuracy of dynamic control content of internet of things equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of cloud collaboration, and in particular to a dynamic control method and system for IoT devices based on cloud collaboration. Background Technology

[0002] With the rapid development of IoT technology and edge computing, cloud-based collaborative architecture has become the mainstream paradigm for realizing dynamic control of large-scale IoT devices, such as industrial robots and automated guided vehicles. In complex industrial operation scenarios, IoT devices not only need to cope with the dynamic changes in their own physical movement, but also need to deal with communication disturbances caused by factors such as uneven wireless network coverage and limited bandwidth.

[0003] In current cloud-based control systems, sensor data and network status data reported by IoT devices are processed in isolation or simply feature-stitched. This fails to construct a unified global state space to extract dynamic feature vectors that include temporal dependencies and spatial coupling. This flattened feature processing approach ignores the deep connections between the physical topology and logical communication of device clusters, resulting in blind spots in the cloud's control over the system's status. At the same time, existing control offset calculations are mostly based on the instantaneous state threshold at the current moment, lacking proactive prediction of future state trajectories in the cloud collaborative space. This affects the accuracy of control offset and leads to low accuracy in the dynamic control content of IoT devices. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a dynamic control method and system for IoT devices based on cloud collaboration.

[0005] This invention provides a dynamic control method for IoT devices based on cloud collaboration, comprising: In the cloud server, multi-source heterogeneous stream data from IoT devices are received, and a global state space of IoT devices is constructed by combining it with a preset spatiotemporal correlation graph. Feature aggregation and dimensionality reduction are performed on the global state space to extract dynamic feature vectors containing temporal dependency characteristics and spatial coupling characteristics. The dynamic feature vector is input into the corresponding cloud collaborative space, and the state prediction trajectory of the IoT device in the future time domain is generated during the feature evolution process. The corresponding control offset is determined by comparing it with the preset safe operation trajectory. When the control offset meets the dynamic collaborative triggering condition, the cloud server solves the optimal control strategy with the joint optimization objectives of delay minimization and control stability rate based on the dynamic feature vector and state evolution trajectory, and determines the collaborative control level that matches the optimal control strategy. The cloud server distributes the optimal control strategy to the corresponding collaborative control node according to the collaborative control hierarchy, and collects observation data after the strategy is executed in real time. It further combines the current load of the cloud server and the current status of the IoT device to determine the dynamic control content of the IoT device.

[0006] This invention provides a dynamic control system for IoT devices based on cloud collaboration. The dynamic control system for IoT devices based on cloud collaboration is applied to the aforementioned dynamic control method for IoT devices based on cloud collaboration. The dynamic control system for IoT devices based on cloud collaboration includes: The dynamic feature module is used to receive multi-source heterogeneous stream data from IoT devices in the cloud server, and construct the global state space of IoT devices by combining it with a preset spatiotemporal correlation graph. It then performs feature aggregation and dimensionality reduction on the global state space to extract dynamic feature vectors containing temporal dependency characteristics and spatial coupling characteristics. The cloud collaboration module is used to input dynamic feature vectors into the corresponding cloud collaboration space, generate the state prediction trajectory of IoT devices in the future time domain during the feature evolution process, and determine the corresponding control offset by comparing with the preset safe operation trajectory. The collaborative control module is used to solve the optimal control strategy with the joint optimization objectives of delay minimization and control stability rate based on dynamic feature vectors and state evolution trajectories when the control offset meets the dynamic collaborative triggering conditions, and to determine the collaborative control level that matches the optimal control strategy. The dynamic control module is used by the cloud server to distribute the optimal control strategy to the corresponding collaborative control node according to the collaborative control hierarchy, and to collect observation data after the strategy is executed in real time. It further combines the current load of the cloud server and the current status of the IoT device to determine the dynamic control content of the IoT device.

[0007] Compared with the prior art, the beneficial effects of the present invention are: (1) In the cloud server, multi-source heterogeneous stream data of IoT devices are received, and the global state space of IoT devices is constructed in combination with the preset spatiotemporal correlation graph. The global state space is subjected to feature aggregation and dimensionality reduction to extract dynamic feature vectors containing temporal dependency characteristics and spatial coupling characteristics. The dynamic feature vectors are input into the corresponding cloud collaborative space, and the state prediction trajectory of IoT devices in the future time domain is generated during the feature evolution process. The corresponding control offset is determined by comparing the preset safe operation trajectory. The introduction of dynamic feature vectors further controls the cloud collaborative space and improves the accuracy of control offset.

[0008] (2) When the control offset meets the dynamic collaborative triggering condition, the cloud server solves the optimal control strategy with the joint optimization objectives of delay minimization and control stability rate based on the dynamic feature vector and state evolution trajectory, and determines the collaborative control level that matches the optimal control strategy. The cloud server distributes the optimal control strategy to the corresponding collaborative control node according to the collaborative control level, and collects the observation data after the strategy is executed in real time. It further combines the current load of the cloud server and the current state of the IoT device to determine the dynamic control content of the IoT device, further controls the collaborative control level, and fully considers the observation data, the current load of the cloud server and the current state of the IoT device, thereby improving the accuracy of the dynamic control content of the IoT device. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating the dynamic control method for IoT devices based on cloud collaboration in an embodiment of the present invention. Figure 2 This is a flowchart illustrating step S11 in the dynamic control method for IoT devices based on cloud collaboration in this embodiment of the invention. Figure 3 This is a flowchart illustrating step S12 in the dynamic control method for IoT devices based on cloud collaboration in this embodiment of the invention. Figure 4 This is a flowchart illustrating step S13 in the dynamic control method for IoT devices based on cloud collaboration in this embodiment of the invention. Figure 5 This is a flowchart illustrating step S14 in the dynamic control method for IoT devices based on cloud collaboration in this embodiment of the invention. Figure 6 This is a schematic diagram of the structure of the dynamic control system for IoT devices based on cloud collaboration in an embodiment of the present invention. Detailed Implementation

[0010] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0011] Please see Figures 1 to 6 A dynamic control method for IoT devices based on cloud collaboration, applied to cloud collaboration scenarios; the dynamic control method for IoT devices based on cloud collaboration includes: Step S11: In the cloud server, receive multi-source heterogeneous stream data from IoT devices, and construct the global state space of IoT devices by combining the preset spatiotemporal correlation graph. Perform feature aggregation and dimensionality reduction on the global state space to extract dynamic feature vectors containing temporal dependency characteristics and spatial coupling characteristics. Step S12: Input the dynamic feature vector into the corresponding cloud collaborative space, and generate the state prediction trajectory of the IoT device in the future time domain during the feature evolution process, and determine the corresponding control offset by comparing it with the preset safe operation trajectory. Step S13: When the control offset meets the dynamic collaborative triggering condition, the cloud server solves the optimal control strategy with the joint optimization objectives of delay minimization and control stability rate based on the dynamic feature vector and state evolution trajectory, and determines the collaborative control level that matches the optimal control strategy. Step S14: The cloud server distributes the optimal control strategy to the corresponding collaborative control node according to the collaborative control hierarchy, and collects observation data after the strategy is executed in real time. It further combines the current load of the cloud server and the current status of the IoT device to determine the dynamic control content of the IoT device.

[0012] refer to Figure 2 In step S11, the specific steps are as follows: S111: During the operation of IoT devices, the cloud server receives multi-source heterogeneous stream data reported by IoT devices in real time based on the streaming computing engine. This multi-source heterogeneous stream data includes multimodal sensing parameters and network status indicators. Combined with graph attention network, a preset spatiotemporal correlation graph is constructed with devices as nodes and physical topology and logical communication links between devices as edges. The multi-source heterogeneous stream data is mapped to the spatiotemporal correlation graph, thereby constructing the global state space of IoT devices. S112: Obtain an autoencoder that introduces a spatiotemporal masking mechanism, perform feature aggregation and dimensionality reduction on the global state space, capture the spatial coupling characteristics of the IoT device cluster through an adaptive graph pooling layer during the aggregation process, and extract the temporal dependency characteristics of node state evolution through a temporal convolutional network, thereby extracting a dynamic feature vector containing temporal dependency characteristics and spatial coupling characteristics.

[0013] In the embodiments of this application, during the operation of IoT devices, the cloud server receives multi-source heterogeneous stream data reported by IoT devices in real time based on the streaming computing engine. The multi-source heterogeneous stream data includes multimodal sensing parameters and network status indicators. Combined with graph attention network, a preset spatiotemporal correlation graph is constructed with devices as nodes and physical topology and logical communication links between devices as edges. The multi-source heterogeneous stream data is mapped to the spatiotemporal correlation graph, thereby constructing the global state space of IoT devices and ensuring the accuracy of the global state space of IoT devices.

[0014] During the operation of IoT devices, cloud servers deploy distributed streaming computing engines (such as Apache Flink) to build real-time data inflow pipelines, continuously monitoring and receiving multi-source heterogeneous streaming data reported by the IoT device cluster. This multi-source heterogeneous streaming data is continuous in the time dimension and includes, in the modal dimension, "multimodal sensing parameters reflecting the physical operating status of the devices, such as temperature, pressure, vibration frequency, and three-dimensional coordinates," as well as "network status indicators reflecting communication quality, such as transmission latency, packet loss rate, and signal-to-noise ratio." Due to differences in the sampling frequency and dimensions of heterogeneous data, the streaming computing engine uses a sliding window mechanism based on event time for timestamp alignment when data flows in, and performs standardization and linear interpolation of missing values, thereby converting the original asynchronous heterogeneous data stream into a unified synchronous multidimensional time series matrix, completing the initial data regularization.

[0015] To address the coupling characteristics of IoT device clusters, a pre-defined spatiotemporal relational graph is constructed, consisting of a two-layer fusion of physical topology edges and logical communication link edges. For physical topology edges, the K-nearest neighbor algorithm or a communication distance threshold is used to determine physical adjacency based on the actual deployment location of the devices in the physical space, generating a physical adjacency matrix. For logical communication link edges, routing tables and data packet interaction frequencies are extracted from network status indicators to calculate the logical dependency strength between devices, generating a logical adjacency matrix. A learnable fusion weight parameter is introduced to linearly weight and fuse the physical adjacency matrix and the logical adjacency matrix, generating a comprehensive adjacency matrix. This matrix reflects both the physical hard connection constraints between devices and the logical soft dependencies of data flow.

[0016] The synchronous multidimensional time series matrix is ​​used as the initial feature of the nodes, and the comprehensive adjacency matrix is ​​used as the topological constraint, both input into the graph attention network (GAT). In each aggregation process of the graph attention network, for any target node, the attention coefficient between it and its neighboring nodes is calculated. Specifically, the node features are mapped to higher dimensions through a shared linear transformation matrix, and the feature similarity between node pairs is calculated using a single-layer feedforward neural network. After processing with the LeakyReLU activation function and Softmax normalization, the dynamic attention weights of neighboring nodes to the target node are obtained. These dynamic attention weights break the static allocation limitation of traditional graph convolution based on topological degree, and can adaptively enhance the feature weights of "critical neighboring nodes, such as abnormal alarm nodes or high-load link nodes" according to the real-time changes of multimodal sensing parameters and network state indicators, thus suppressing noise interference from redundant nodes.

[0017] Through multi-layer stacking and feature aggregation of graph attention networks, the feature vector of each node not only integrates its own multimodal sensing and network state, but also aggregates the spatiotemporal context information of physically and logically adjacent nodes within a multi-hop range. By concatenating and non-linearly activating the aggregated feature vectors of all nodes in the graph dimension, a spatiotemporal correlation graph containing the complete operational status of the device cluster is generated. At this point, the node feature matrix and topological edge relationships of this spatiotemporal correlation graph together constitute the global state space of the IoT devices. This global state space is no longer an isolated single device dataset, but a dynamic manifold representation that deeply integrates spatial physical coupling and real-time communication status.

[0018] Specifically, in a cloud server collaborative control scenario in a smart factory, there are 50 automated guided vehicles (AGVs) working collaboratively. In this case, AGVs are IoT devices. During operation, AGVs need to avoid obstacles and collaboratively transport materials, and are affected by uneven wireless network coverage within the factory. The 50 AGVs continuously report data to the cloud at a frequency of 10Hz. Multimodal sensing parameters include the AGV's current 3D coordinates (X, Y, Z), battery level, and motor speed; network status indicators include communication latency and packet loss rate between the AGV and the AP (wireless access point). The cloud-based Flink engine receives this heterogeneous data stream. Due to clock drift in AGV reporting, Flink uses a 500ms sliding window for timestamp alignment and normalizes the coordinates and speed data with different dimensions, forming a real-time synchronized feature matrix of shape [50, 7], representing 50 devices and 7 feature dimensions.

[0019] The edges of the spatiotemporal relational graph are constructed in the cloud: At the physical topology level, the Euclidean distance between each pair of AGVs is calculated. If the distance is less than 5 meters, indicating a risk of collision or potential collaborative transport, a physical edge is established. At the logical communication link level, if two AGVs belong to the same multicast group or have direct task signaling interaction, a logical edge is established. These two aspects are then merged. For example, if AGV_01 and AGV_02 are 3 meters apart and are engaged in collaborative transport, their merged edge has a very high weight; while if AGV_01 and AGV_50 are 200 meters apart and have no communication, their edge weight is 0. This generates a 50×50 sparse merged adjacency matrix.

[0020] The feature matrix and adjacency matrix of [50, 7] are input into the GAT network. If AGV_01 experiences an emergency stop, its features change drastically. When GAT aggregates features for neighboring AGV_02 and AGV_03, it automatically calculates and assigns a very high attention weight (e.g., 0.8) to AGV_01 through a dynamic attention mechanism, while assigning lower weights (e.g., 0.2) to other normally functioning neighbors. This results in the node features of AGV_02 and AGV_03 being strongly loaded with the abnormal state of their neighbor AGV_01.

[0021] After two layers of GAT feature aggregation, the cloud obtained 50 new high-dimensional feature vectors for the AGVs. These vectors not only contain the AGV's own power and coordinates, but also deeply integrate the risk of sudden stops of AGVs within a 5-meter radius and the congestion status of the communication link. These 50 high-dimensional feature vectors, together with the fused adjacency matrix, constitute the "global state space" of the AGV cluster at the current moment. In subsequent steps, the cloud uses this global state space to perceive the spatial coupling impact of AGV_01's fault on the surrounding AGVs, thereby predicting the trajectory in advance and issuing avoidance control strategies, rather than passively responding only after AGV_02's own sensors detect obstacles.

[0022] Furthermore, an autoencoder with a spatiotemporal masking mechanism is obtained to perform feature aggregation and dimensionality reduction on the global state space. During the aggregation process, the spatial coupling characteristics of the IoT device cluster are captured through an adaptive graph pooling layer, and the temporal dependency characteristics of node state evolution are extracted through a temporal convolutional network, thereby extracting a dynamic feature vector containing temporal dependency characteristics and spatial coupling characteristics.

[0023] At this point, a pre-trained autoencoder architecture is acquired. Before the encoder performs feature aggregation and dimensionality reduction on the global state space, a spatiotemporal masking mechanism is introduced to randomly mask the input data. In the spatial dimension, based on the set spatial masking ratio, the feature vectors of some nodes in the spatiotemporal correlation graph are randomly selected and set to zero to simulate the data upload loss scenario caused by wireless signal attenuation or failure of IoT devices. In the temporal dimension, based on the set temporal masking span, the node features of continuous time steps are subjected to fragment-level masking processing to simulate the time-series data gaps caused by the instantaneous interruption of the communication link. The damaged global state space is generated through this spatiotemporal masking operation, forcing the autoencoder to not only memorize the input data during the feature dimensionality reduction process, but also to learn the potential spatiotemporal correlations between nodes in order to accurately reconstruct the feature representations that are masked.

[0024] The masked global state space is input into the encoder of the autoencoder. The high-dimensional graph structure is hierarchically compressed through an adaptive graph pooling layer to capture the spatial coupling characteristics of the IoT device cluster and eliminate redundant topology. In the adaptive graph pooling layer, the learnable projection score of each node in the graph is calculated. This score comprehensively reflects the importance of the node's own features and the tightness of its topological connections with neighboring nodes. The nodes are sorted in descending order according to the projection scores, and the nodes with the highest ranking are truncated and a set proportion is retained. At the same time, a corresponding coarsened adjacency matrix is ​​generated, and low-scoring nodes and their connecting edges are discarded. Through this operation, highly coupled device nodes with close physical distances and highly consistent feature evolution are dynamically clustered and mapped as super nodes. Thus, while preserving the global key spatial topological dependencies, the size of the graph is greatly compressed, and a compact spatial coupling feature matrix is ​​extracted.

[0025] After completing the adaptive graph pooling compression in the spatial dimension, a temporal convolutional network is used to extract the temporal dependency characteristics of node state evolution for the compressed supernode feature sequence. This temporal convolutional network adopts an architecture composed of multiple layers of dilated causal convolutions. The causal convolution mechanism ensures that the feature extraction process only depends on the input of the historical time and the current time, eliminating the leakage of future information and meeting the real-time constraints of dynamic control. The dilated convolution mechanism expands the receptive field of the convolution kernel by increasing the dilation factor exponentially, enabling the network to capture multi-scale temporal dependency characteristics from short-period high-frequency oscillations to long-period slow-changing trends under the condition of constant parameters. Furthermore, residual connections and weight normalization operations are introduced between layers to alleviate the gradient vanishing problem during deep network training, and finally output a temporal dependency feature matrix that has both long and short-term historical memories.

[0026] The spatial coupling feature matrix output by the adaptive graph pooling layer and the temporal dependency feature matrix output by the temporal convolutional network are concatenated and fused along the feature channel dimension. The fused feature tensor is then nonlinearly mapped and dimensionality reduced through a bottleneck structure composed of fully connected layers, compressing the high-dimensional tensor into a low-dimensional latent variable with extremely high information density. This latent variable is the extracted dynamic feature vector containing temporal dependency and spatial coupling characteristics. In a very low-dimensional space, this dynamic feature vector not only encodes the coupling relationship between the physical and logical spaces of the device cluster, but also embeds its dynamic law of evolution over time.

[0027] Specifically, the cloud acquires a global state space constructed from 50 AGVs, including features such as coordinates, speed, and communication latency for each AGV. Before inputting the data into the autoencoder, the cloud introduces a spatiotemporal masking mechanism: spatially, the current feature vectors of 5 AGVs (e.g., AGV_05, AGV_12, etc.) are randomly set to zero, simulating data loss caused by these AGVs traveling into a wireless signal blind spot in the factory; temporally, the data from the past 3 sampling periods of certain AGVs are masked, simulating brief communication interruptions. The autoencoder's task in subsequent compression is to predict the state of the 5 masked AGVs based on the data and topology association of the remaining 45 AGVs.

[0028] Entering the encoder's pooling phase, the adaptive graph pooling layer begins processing the topology of these 50 AGVs. Since AGV_01, AGV_02, and AGV_03 are currently jointly performing a collaborative material handling task, their physical distance is less than 2 meters and they communicate frequently logically. Their feature evolution in the graph space is highly consistent, and their calculated projection scores are highly correlated. The pooling layer dynamically merges these three highly coupled nodes into a single supernode of a "handling collaboration group," and the edge connections are coarsened accordingly. Through this operation, the massive topology of 50 nodes is compressed into, for example, 20 supernodes. Each supernode represents a spatially highly coupled local part of the AGV cluster, accurately capturing the spatial coupling characteristics brought about by collaborative handling, while eliminating redundant edges from AGVs far from the cluster and operating independently.

[0029] For the 20 pooled super nodes, the cloud uses a Temporal Convolutional Network (TCN) to extract their temporal features. For example, for the "transportation coordination group" super node mentioned above, its speed feature in the past 10 seconds shows a short-period fluctuation of "acceleration-uniform speed-slight deceleration", while its battery power shows a long-period slow decreasing trend. Through dilated causal convolution, the TCN captures the high-frequency obstacle avoidance feature of slight deceleration with a small dilation factor and the low-slow trend feature of battery power decrease with a large dilation factor, and strictly follows the causal law, thereby extracting the temporal dependency of the super node's state evolution.

[0030] The cloud concatenates the "spatial compression features" of the "transportation coordination group" supernode, which characterize the coordination tightness and relative position constraints of the three AGVs within the group, with the "temporal evolution features" extracted by TCN, which characterize the overall acceleration / deceleration inertia and power consumption trends of the group. After dimensionality reduction through a fully connected layer, a dynamic feature vector with a length of only 64 is generated. At this point, the massive amount of data, originally containing 50 AGVs, multi-dimensional heterogeneous parameters, and complex spatiotemporal relationships, is extremely condensed into these 64 floating-point numbers. These 64 floating-point numbers not only know how the AGV group is currently moving (temporal dependence) but also who is moving together with whom (spatial coupling). When the cloud subsequently inputs this vector into the coordination space for trajectory pre-simulation, the computational complexity decreases exponentially, achieving a foundation for highly real-time dynamic control.

[0031] refer to Figure 3 In step S12, the specific steps are as follows: S121: Obtain the cloud collaboration space deployed in the cloud, input the dynamic feature vector into the cloud collaboration space, determine the state evolution event based on the fusion architecture of long short-term memory network and graph neural network, and generate the multi-step state pre-trajectory trajectory of IoT device in the future time domain by deduce the nonlinear state transition of dynamic feature vector in the state evolution event. S122: Obtain the preset safe operation trajectory, which is generated by the physical boundary and operating condition constraints of the IoT device. Dynamically align and compare the state pre-simulation trajectory and the safe operation trajectory in the same space. Determine the weighted sum of the distance and trajectory slope deviation between the state pre-simulation trajectory and the safe operation trajectory in the same space. Quantize the weighted sum into a control offset degree that characterizes the degree of deviation of the system from the safe state.

[0032] In the embodiments of this application, a cloud-based collaborative space deployed in the cloud is obtained. Dynamic feature vectors are input into the cloud-based collaborative space. State evolution events are determined based on the fusion architecture of long short-term memory network and graph neural network. In the state evolution events, the nonlinear state transition of dynamic feature vectors is deduced to generate multi-step state prediction trajectories of IoT devices in the future time domain. This is compatible with the overall consideration of the fusion architecture of long short-term memory network and graph neural network, ensuring the accuracy of state evolution events.

[0033] At this point, a cloud-based collaborative space deployed in the cloud is acquired. This cloud-based collaborative space is a high-fidelity mapping container for IoT device clusters on the digital side, embedding device kinematic constraints and communication topology evolution rules. The dynamic feature vectors extracted in the previous steps, containing temporal dependency and spatial coupling characteristics, are input into this cloud-based collaborative space. During the input mapping stage, the dynamic feature vectors are re-expanded into a hidden graph structure tensor representing the holographic state of the device cluster at the current moment through the state initialization module. This tensor contains the hidden state characteristics of each node and the reconstructed adjacency relationship, serving as the initial boundary conditions and dynamic starting field for triggering the cloud-based collaborative space to perform future state inference.

[0034] Within the cloud-based collaborative space, a fusion architecture of Long Short-Term Memory (LSTM) and Graph Neural Network (GNN) is constructed to determine state evolution events. The temporal gating mechanism of the LSTM is embedded into the message passing process of the GNN, forming a spatiotemporally synchronized computational unit. In each inference step, the GNN layer is responsible for aggregating the spatial coupling features of adjacent nodes in the topological dimension, generating a spatial context vector. This spatial context vector, along with the node's own historical hidden state, is input to the forget gate, input gate, and output gate of the LSTM. The gating mechanism adaptively filters and retains long-term, slowly changing spatial constraint memories while updating short-term, rapidly changing local response features. A complete forward propagation of the spatiotemporally synchronized computational unit is defined as a state evolution event, representing a nonlinear state transition paradigm of the device cluster under a small timescale, governed by both physical laws and logical coupling.

[0035] Based on deterministic state evolution events, nonlinear state transition deduction is performed on dynamic feature vectors. During the deduction process, the implicit graph structure tensor of the previous moment is used as input, and the graph feature transition increment of the current moment is calculated through the fusion architecture. This transition increment is superimposed on the implicit state of the previous moment to realize the recursive calculation of the nonlinear state transition function. In this process, the weight parameters of the fusion architecture implicitly encode nonlinear physical and network constraints such as device dynamic friction and communication link congestion backoff, so that the state transition process is no longer a simple linear extrapolation, but accurately approximates the nonlinear evolution manifold of the real physical system under complex working conditions, thereby deducing the real motion trajectory of the cluster state in the potential phase space.

[0036] The nonlinear state transition deduction process is unfolded in an autoregressive manner along the time axis in multiple iterative steps. In each step of the deduction, the implicit graph structure tensor generated in the previous step is used as a new dynamic input, and the state evolution event is repeatedly executed to obtain the state prediction value for the next time step. By continuously executing the autoregressive iteration for a set number of steps, the cloud collaborative space sequentially outputs the predicted state tensor sequence corresponding to each discrete time point in the future time domain. The tensor sequence is demapped in the node dimension and feature dimension to finally generate the multi-step state pre-deduction trajectory of IoT devices in the future time domain. This trajectory covers the multimodal sensor parameter prediction values ​​and network state index prediction values ​​of each device in the cluster at future times, providing a continuous and forward-looking digital twin observation sequence for the subsequent calculation of control offset.

[0037] Specifically, the cloud-deployed collaborative space is essentially a digital sandbox built for these 50 AGVs. The highly condensed 64-dimensional dynamic feature vector output from the previous steps is input into this sandbox. At the moment of input, this vector is decompressed and initialized by the module, and then expanded back into 50 nodes and their corresponding latent feature matrices and adjacency matrices. This is like instantly restoring the precise position, speed, and communication link status of these 50 AGVs on the digital sandbox.

[0038] The cloud-based collaborative space does not employ the traditional sequential logic of calculating time first and then space. Instead, it uses an architecture that integrates LSTM and GNN to determine state evolution events. Taking AGV_01 as an example, when deducing its state in the next second, the GNN layer receives spatial context information (such as their relative speed and distance) from its surrounding physical and logical neighbors (e.g., AGV_02 and AGV_03, which are currently cooperating in transport). The LSTM gating mechanism intervenes: the forget gate decides to discard the slight communication packet loss interference from the previous second, while the input gate writes the significant spatial correlation signal of AGV_02 and AGV_03 decelerating into AGV_01's long-term memory. The output gate finally synthesizes the new hidden state of AGV_01 under the influence of its neighbors. This spatiotemporal fusion deduction step is a state evolution event, which accurately reflects the linkage response mechanism of the AGV cluster.

[0039] Based on the aforementioned state evolution events, the cloud begins to perform nonlinear state transition simulations on the AGV cluster. Traditional linear simulations might assume that AGV_01 will maintain its current constant speed, but the nonlinear transition function of the fusion architecture implicitly includes the AGV's "dynamic constraints, such as braking distance during a full-load emergency stop" and "network backoff constraints, such as a sudden increase in communication latency caused by channel contention due to the convergence of multiple AGVs." Therefore, when AGV_01 receives an emergency stop command, the simulation process accurately presents its speed as a nonlinear decrease followed by a slow approach to zero gliding curve. Simultaneously, its network latency also decreases nonlinearly because it no longer needs to transmit real-time speed commands. This step of the simulation approximates the complex nonlinear characteristics of the real physical world.

[0040] The cloud-based collaborative space continuously rolls the aforementioned nonlinear state transition deduction forward in an autoregressive manner. Using the AGV's latent state deduced from the previous second as new input, the LSTM-GNN fusion architecture is triggered again to predict the state for the next second. Assuming a future time domain of 5 seconds and a sampling interval of 0.5 seconds, this process iterates 10 times. The cloud acquires a complete set of predicted data sequences, including coordinates, speeds, and communication delays, for 50 AGVs at 10 discrete time points from the current moment to the next 5 seconds. This is the multi-step state prediction trajectory, enabling the cloud controller to anticipate potential congestion or collisions within the next 5 seconds and issue dynamic control commands in advance.

[0041] Furthermore, a preset safe operating trajectory is obtained, which is generated by the physical boundaries and operating condition constraints of the IoT device. The state pre-simulation trajectory and the safe operating trajectory are dynamically aligned and compared in the same space. The weighted sum of the distance and trajectory slope deviation between the state pre-simulation trajectory and the safe operating trajectory in the same space is determined. The weighted sum is quantified as the control offset degree characterizing the degree of deviation of the system from the safe state. This takes into account the overall consideration of the state pre-simulation trajectory and the safe operating trajectory in the same space. At the same time, dynamic feature vectors are introduced to further control the cloud collaborative space and improve the accuracy of the control offset degree.

[0042] At this point, a preset safe operating trajectory is obtained. This safe operating trajectory is not a single scalar threshold, but a multi-dimensional spatiotemporal envelope channel jointly defined by the physical boundaries and operating condition constraints of the IoT devices. At this point, the physical boundaries of the devices under the mechanical kinematic limits are extracted. At the same time, the operating condition constraints set by the process flow and task scheduling, such as the speed limit requirements of specific work areas, the collision avoidance safety distance threshold, and the safety upper limit of network communication latency, are combined to construct a dynamic envelope surface in the state space with the nominal planned trajectory as the center and the physical and operating condition constraints as the upper and lower bounds. This envelope surface extends along the time axis to form a safe operating trajectory channel with time-varying bandwidth, which serves as an absolute reference benchmark for measuring whether the future state of the device cluster is safe and compliant.

[0043] The state pre-simulation trajectory generated in the previous steps and the safe operation trajectory are loaded into the same high-dimensional state space for dynamic alignment and comparison. Since the evolution of the state pre-simulation trajectory on the time axis may cause elastic offset on the time axis from the nominal safe operation trajectory due to the lag or advance of the device's dynamic response, a dynamic time warping algorithm is introduced. During the comparison process, by constructing a cumulative distance matrix, an optimal nonlinear mapping path that minimizes the feature distance between the pre-simulation trajectory and the safe trajectory is found, thereby overcoming the rigid misalignment between the two on the time axis and ensuring that subsequent distance calculations are performed under the same physical phase and logical stage of device operation, achieving accurate alignment of the trajectory in the spatial and temporal dimensions.

[0044] After completing the dynamic alignment, for each time step in the future time domain, the spatial relative distance and trajectory slope deviation between the pre-simulated trajectory and the safe operating trajectory are decoupled and calculated. For the spatial relative distance, the Euclidean distance from the state point of the pre-simulated trajectory to the boundary of the safe envelope channel is calculated. If the state point is located inside the envelope channel, the distance is negative or zero; if it penetrates the envelope boundary, it is positive, in order to quantify the depth of the system's overshoot. For the trajectory slope deviation, the first-order difference vector of the pre-simulated trajectory at the current time step is calculated to represent the instantaneous change trend. The cosine distance or vector L2 norm difference between the pre-simulated trajectory and the first-order difference vector of the safe operating trajectory at the corresponding time step is calculated to represent the degree to which the dynamic evolution trend of the system deviates from the safe evolution trend, i.e., at what rate of instability the system is leaving the safe channel.

[0045] The spatial relative distance and trajectory slope deviation calculated at the same time step are adaptively weighted and summed, and the result is accumulated or taken as an extreme value along the time axis, and finally quantified as the control offset that characterizes the degree of deviation of the system from the safe state. In the weighting process, the weight coefficients are dynamically adjusted according to the severity of the operating conditions: when the physical boundary margin is extremely small, the spatial relative distance is given a higher weight to punish the risk of immediate over-limit; when the equipment is in a high-speed dynamic changing condition, the trajectory slope deviation is given a higher weight to punish the risk of trend instability. By performing time-domain maximum pooling or decay summation on the weighted sum of future multi-step steps, the control offset in scalar form is output. This value directly maps the probability and severity of the system going out of control in the future time domain, and serves as the core criterion for triggering dynamic cooperative control.

[0046] Regarding control offset, to ensure that the control offset can be accurately calculated and used as a reliable criterion for dynamic cooperative control, this invention provides explicit quantitative calibration of the weighting coefficients in the weighted summation process and the offset threshold used to trigger control, as follows: The weighting coefficients used for spatial relative distance and trajectory slope deviation in calculating control offset are not fixed values, but dynamically determined based on the real-time operating conditions of the IoT device. Specifically, an adaptive weighting factor is defined, ranging from 0 to 1, which is jointly determined by the device's instantaneous motion state and the urgency of the task. When the device's state-previewed trajectory identifies it as being in a rapidly changing condition such as high-speed movement, sharp turns, or performing high-precision docking, the system sets the weighting coefficient for trajectory slope deviation to a higher first value, such as 0.7, while simultaneously setting the weighting coefficient for spatial relative distance to a complementary second value, such as 0.3, to prioritize suppressing the system's tendency to become unstable.

[0047] Conversely, when the equipment is in low-speed cruising, stationary waiting, or performing rough handling, the system will reduce its sensitivity to trend deviations, adjust the weighting coefficient of trajectory slope deviation to a lower third value, such as 0.2, and correspondingly increase the weighting coefficient of spatial relative distance to 0.8, so as to focus on monitoring whether the equipment position exceeds the physical safety boundary. Under normal operating conditions that are not boundary or extreme, the two weighting coefficients are set to be equal by default, such as both being 0.5, so as to achieve a balanced assessment of instantaneous over-limit and trend deviation.

[0048] The offset threshold used to determine whether the dynamic collaborative triggering conditions are met is not a static constant, but is dynamically calibrated and maintained online based on the statistical characteristics of the IoT device cluster during historical steady-state operation. When the system is initially put into operation or during regular maintenance, the cloud server will collect normal operation data within a relatively long historical window and calculate the control offset value corresponding to each control cycle within that period, forming a benchmark sample set that characterizes the healthy operating status of the system.

[0049] The system performs statistical analysis based on this set, calculating its mean and standard deviation. The preset offset threshold is dynamically determined based on the sum of the mean and three times the standard deviation of this benchmark sample set. This setting method statistically covers 99.7% of the fluctuation range under normal operating conditions. In actual operation, whenever the real-time calculated value of the control offset exceeds this dynamic threshold, it is determined that it deviates from the statistical upper limit of the historical steady-state law, satisfying the dynamic collaborative triggering condition. To adapt to long-term slowly changing factors such as equipment aging or gradual changes in the operating environment, the cloud server also uses a sliding time window to periodically recalculate this statistical threshold using recent normal operating data, thereby ensuring that the triggering criteria always match the current actual health status of the system and avoiding false triggering or missed triggering caused by using a fixed threshold.

[0050] Specifically, the cloud sets a safe operating trajectory for AGV_01, which performs the handling task. This trajectory is not simply a "route," but a three-dimensional spatiotemporal envelope channel: at the physical boundary, the width of the channel is formed by the physical dimensions of AGV_01 plus a 0.5-meter collision avoidance safety distance, and the height is truncated by the maximum speed limit of 1.5 m / s; in terms of operating condition constraints, when AGV_01 enters the blind zone of an intersection, the channel bandwidth will dynamically narrow, and the safe upper limit for network latency is set at 100 ms. This envelope channel with flexible width and multi-dimensional constraints is the safe operating trajectory of AGV_01; as long as its future state remains within this channel, it is considered safe.

[0051] The cloud compares the previously deduced state prediction trajectory of AGV_01 for the next 5 seconds with the safety envelope channel in the same state space. Assume that due to wireless network jitter, AGV_01's instruction execution experiences a 0.3-second delay, causing its prediction trajectory to lag behind the nominal center line of the safety trajectory on the time axis. A direct comparison would show a significant positional discrepancy; however, through Dynamic Time Warping (DTW) algorithm, the cloud identifies that the "acceleration segment" of the prediction trajectory and the "acceleration segment" of the safety trajectory are physically and logically in phase, only elastically misaligned in time. This stretches and compresses the two on the time axis, achieving precise alignment of the physical motion phases and ensuring that subsequent calculations represent the "deviation under the same motion stage."

[0052] After alignment, the cloud decouples the calculation of the deviation in the two dimensions. Assuming that in the 3-second rehearsal, AGV_01 deviates 0.6 meters to the right to avoid a sudden obstacle, penetrating the 0.5-meter safety envelope boundary, the calculated relative spatial distance is +0.1 meters, with a positive value indicating the depth of the deviation. Simultaneously, the rehearsal trajectory shows AGV_01 rapidly turning its steering wheel back to center at an acceleration of 2.0 m / s², while the safety trajectory at this phase requires uniform straight-line movement. The significant difference in the rate of change of their velocity vectors results in a large deviation in the calculated trajectory slope. This indicates that the AGV has not only crossed the boundary but is also experiencing a severe dynamic instability.

[0053] The cloud-based system performs a weighted sum of the distance and slope. Since AGV_01 is currently in a high-speed, variable-condition driving phase, the system adaptively assigns a higher weight to the trajectory slope deviation (e.g., 0.7) and a lower weight to the spatial distance (e.g., 0.3), because drastic trend deviations are often more likely to cause rollovers or chain collisions than instantaneous positional deviations. After the weighted calculation, the system extracts the largest weighted sum within the next 5 seconds to quantify the control offset. For example, a result of 0.85 far exceeds the safety threshold of 0.3. This 0.85 control offset serves as a highly condensed alarm signal, precisely indicating to the cloud that AGV_01 has not only deviated from its safe position but its future movement trend is rapidly deteriorating, necessitating immediate triggering of high-priority dynamic collaborative control rather than waiting for it to recover on its own.

[0054] refer to Figure 4 In step S13, the specific steps are as follows: S131: Real-time monitoring of control deviation. When the control deviation exceeds the preset deviation threshold, it is determined that the dynamic collaborative triggering condition is met. The cloud server constructs a Markov decision process based on dynamic feature vectors and state evolution trajectories, transforming the control strategy solution into a multi-objective optimization problem with the joint optimization objectives of minimizing end-to-end communication delay and maximizing the dynamic response control stability rate of the system. Further, it combines reinforcement learning mechanism to explore and solve the problem, and selects non-dominated solution sets to determine the optimal control strategy. S132: Extract the frequency and criticality features of control actions in the optimal control strategy, combine them with the topological hop count of the cloud-edge collaborative space, and use a hierarchical classifier to determine the collaborative control level that matches the optimal control strategy. The collaborative control level includes global overall control in the cloud, local control in the edge gateway, and local self-determination control of the terminal device.

[0055] In the embodiments of this application, the control offset is monitored in real time. When the control offset exceeds a preset offset threshold, it is determined that the dynamic collaborative triggering condition is met. The cloud server constructs a Markov decision process based on dynamic feature vectors and state evolution trajectories, transforming the solution of the control strategy into a multi-objective optimization problem with the joint optimization objectives of minimizing end-to-end communication latency and maximizing the dynamic response control stability rate of the system. Further, the solution is explored by combining reinforcement learning mechanism, and non-dominated solution sets are selected to determine the optimal control strategy. This approach takes into account the overall consideration of non-dominated solution sets and ensures the accuracy of the optimal control strategy.

[0056] At this point, a high-frequency polling and monitoring mechanism is established in the cloud server to obtain the control offset calculated and output by the previous steps in real time. The control offset is compared with a preset offset threshold in real time. This offset threshold is dynamically calibrated based on the historical operating baseline and fault tolerance capability of the IoT device cluster. When the control offset is less than or equal to the preset offset threshold, the system is determined to be in a safe convergence state, and the current monitoring state is maintained. When the control offset exceeds the preset offset threshold, it indicates that the future state of the system will break through the physical boundary or operating condition constraints. At this time, it is immediately determined that the dynamic collaborative triggering condition is met, the passive monitoring mode is forcibly interrupted, and the cloud-based active intervention and control strategy solution process is triggered.

[0057] When the dynamic collaboration triggering conditions are met, the cloud server constructs a Markov Decision Process (MDP) based on the dynamic feature vector and state evolution trajectory of the current time window. At this time, the dynamic feature vector containing spatiotemporal dependencies is spliced ​​and fused with the future multi-step state pre-simulation trajectory to define the state space of the MDP, so as to comprehensively represent the current physical condition and future evolution trend of the equipment cluster. The "set of control commands that can be issued from the cloud, such as speed regulation rate, heading deflection angle, communication link switching commands, etc." is defined as the action space. The state transition mapping under the action drive is defined as the state transition probability, which is implicitly represented by the inference model of the underlying cloud collaboration space. Thus, the continuous dynamic control problem is strictly transformed into a Markov process of sequential decision-making.

[0058] Within the constructed MDP framework, the solution of the control strategy is transformed into a multi-objective optimization problem, defining a joint optimization objective function. The first optimization objective is to minimize end-to-end communication latency, which penalizes the time loss during the process of generating control commands from the cloud, transmitting them through the network to the edge nodes, and waking up the actuators. Its value is jointly determined by network state indicators and the collaborative architecture level. The second optimization objective is to maximize the stability rate of the system's dynamic response control, which rewards the speed at which the system state pre-simulation trajectory converges into the safe operating trajectory envelope after the control action is injected and the degree of overshoot suppression. Its value is determined by physical boundary constraints and the evolution slope of the dynamic feature vector. The two objectives often have a competitive relationship of mutual constraint in resource-constrained and dynamically evolving scenarios, forming a nonlinear multi-objective optimization space.

[0059] A multi-objective reinforcement learning mechanism is introduced to explore and solve the multi-objective optimization space. A deep neural network is used as a function approximator for the policy network and the value network. The policy network takes the state space of the MDP as input and outputs the probability distribution in the action space. During training and online inference, the agent interacts with the environment based on the current policy to generate a state-action-reward trajectory. To address the multi-objective characteristics, a dual-channel reward function is designed to calculate the delay penalty reward and the stability improvement reward, respectively. The dual-channel reward signals are simultaneously fed back to update the parameters of the value network and the policy network, guiding the agent to gradually approach the Pareto front of multi-objective optimization during the exploration process, rather than getting trapped in the local optimum of single-objective optimization.

[0060] For the dual-channel reward function, which consists of a delay penalty channel and a stability improvement channel, the outputs of both channels are normalized to a uniform numerical range to facilitate the stability of subsequent neural network training. In the delay penalty channel, when the agent generates a control policy under the current state and action, the system measures in real time the total time from policy generation to the policy reaching the target actuator. This total time includes the sum of network transmission latency and cloud queue waiting latency.

[0061] If the total time exceeds the preset safe delay limit, the reward function outputs a negative penalty, and the penalty magnitude is proportional to the length of time exceeding the limit; if the total time is below the delay limit, a small positive reward is output. In the stability improvement channel, the system comprehensively evaluates the speed at which the system state simulation trajectory converges to the safe operating trajectory envelope after the control strategy is executed, as well as the overshoot.

[0062] At this point, if the pre-trained trajectory smoothly enters the safe envelope within the preset time window without repeatedly crossing the boundary, a positive reward is given; if the trajectory exhibits oscillating convergence or continuously exceeds the envelope boundary, a corresponding negative penalty is given based on the magnitude and duration of the deviation. The reward values ​​of the two channels jointly drive the parameter updates of the policy network in the form of a weighted sum during subsequent training, where the initial values ​​of the weights are set at the factory according to the preferences of the application scenario and remain unchanged during training.

[0063] In the continuous action distribution of the reinforcement learning policy network output, a non-dominated ranking algorithm is used to filter the sampled candidate policy set. At this time, for any two policies in the candidate set, if one policy is not inferior to the other in both latency and stability, and is strictly superior to the other policy in at least one of the objectives, then the former is said to be non-dominated to the latter. According to this rule, all dominated inferior solutions are eliminated, and the set of non-dominated solutions that do not dominate each other is retained. This set of non-dominated solutions constitutes the Pareto optimal frontier. Combining the current network load status of the cloud server and the urgency of the real-time operating conditions of the device cluster, an adaptive preference weight is introduced to make a final decision on the non-dominated solution set, and the single policy that best fits the current system status is selected. This policy is determined as the optimal control policy and output to the subsequent distribution module.

[0064] For the non-dominated solution set, for any two candidate strategies in this set, the system calculates their evaluation values ​​on two objectives: "end-to-end communication latency" and "system dynamic response control stability rate." If the first candidate strategy performs no worse than the second candidate strategy on both objectives, and performs significantly better than the second candidate strategy on at least one objective, then the first candidate strategy is determined to be non-dominated. Based on this determination rule, the system traverses the entire candidate strategy set, eliminating all strategies that are dominated by at least one other candidate strategy. The final set of strategies retained is the non-dominated solution set. In actual decision-making, the system additionally obtains the average load rate of the cloud server and the average trajectory deviation of the device cluster at the current moment as preference references, and selects a strategy from the non-dominated solution set that achieves the best engineering balance between the two objectives as the final output and the optimal control strategy to be issued and executed.

[0065] Specifically, the cloud platform monitors the control offset of 50 AGVs in real time. Normally, when AGVs slightly deviate from their routes, the offset is below 0.2, which is lower than the preset threshold of 0.5, and the cloud does not intervene. Suddenly, during collaborative transport, AGV_01 and AGV_02 experience a sudden increase in communication delay due to interference from the factory's wireless signals operating on the same frequency. This causes the trajectory simulation of the two vehicles to indicate an impending collision, and the calculated control offset spikes to 1.8, far exceeding the threshold of 0.5. The cloud platform immediately determines that the dynamic collaborative trigger condition has been met and switches from the normal monitoring mode to the emergency control strategy solution mode.

[0066] The cloud extracts the current "dynamic feature vectors" of AGV_01 and AGV_02, including the battery voltage, motor temperature, and relative distance of the two vehicles, as well as their "state prediction trajectory for the next 5 seconds, showing that the two vehicles will overlap and collide in 3 seconds," and splices them into the state space of the MDP. The action space is defined as the commands that can be issued by the cloud, such as "AGV_01 slows down by 50%", "AGV_02 turns left by 15 degrees", or "switch AGV_01 to the 5G backup link", etc. In this way, the physical evolution process of the two vehicles about to collide is formalized into a Markov decision process that can be intervened in by actions.

[0067] To mitigate collision risks, the cloud-based system employs two interdependent optimization objectives: first, minimizing end-to-end communication latency, as collision avoidance commands must be issued extremely quickly (e.g., 10ms via edge gateway, 80ms via cloud); and second, maximizing control stability, ensuring smooth convergence without causing large materials to overturn or the AGV to tip over during sudden braking or sharp turns. However, rapidly issuing simple emergency stop commands via edge gateways often fails to guarantee material stability, while calculating extremely smooth obstacle avoidance trajectories in the cloud is constrained by the time required for command delivery. These two objectives strongly compete, constituting a multi-objective optimization problem.

[0068] A pre-trained multi-objective reinforcement learning model is launched in the cloud for exploratory solutions. The agent attempts various combinations of actions in the digital twin environment: some actions cause AGV_01 to brake suddenly with extremely low latency, but result in severe material swaying and a deduction in stability score; other actions allow both vehicles to slowly circle around, achieving perfect stability, but the computation and distribution time is too long, resulting in a delay deduction. A dual-channel reward function evaluates the delay penalty and stability gain of each attempt in real time. The neural network continuously adjusts its parameters, gradually learning to find a balance point and exploring a series of candidate control strategies with different degrees of compromise.

[0069] The reinforcement learning model ultimately outputs a batch of candidate strategies, which are then filtered by the cloud using non-dominated ranking. For example, "Strategy A: 15ms latency, 75% stability" and "Strategy B: 40ms latency, 95% stability"—no strategy can beat them in both latency and stability—forming a non-dominated solution set. At this point, the cloud makes a final decision based on the current operating conditions: since the pre-simulated trajectory indicates a collision will occur in 3 seconds, there is still time to spare, and the precision equipment being handled is extremely sensitive to vibration, the cloud adaptively increases the preference weight for "stability," ultimately selecting "Strategy B: 40ms latency, 95% stability" from the solution set as the optimal control strategy. This ensures that, without material tipping, instructions can safely reach the AGV for execution before a collision.

[0070] Furthermore, the frequency and criticality features of control actions in the optimal control strategy are extracted. Combined with the topological hop count of the cloud-edge collaborative space, a hierarchical classifier is used to determine the collaborative control level that matches the optimal control strategy. The collaborative control level includes cloud-based global overall control, edge gateway local control, and terminal device local self-determination control.

[0071] At this point, after determining the optimal control strategy, the action sequence of the strategy in the time domain is deconstructed to extract the frequency and criticality features of the control actions. For the frequency feature, the number of control commands occurring and the time interval between adjacent actions within the set future control time domain are statistically analyzed to calculate the action switching frequency. If the action switching frequency is higher than a preset high-frequency threshold, it indicates that the control strategy requires high-frequency fine-tuning of the equipment, placing extremely high demands on the real-time throughput and low latency characteristics of the communication link. For the criticality feature, the amplitude and range of the control action are extracted. If the action amplitude touches the physical output extreme value of the actuator, or if the execution of the action will cause a drastic phase change in the spatial coupling state of multiple devices, a high criticality weight is assigned, indicating that the control action is a core intervention that determines the system's safety baseline, with extremely low fault tolerance. This generates a two-dimensional feature vector representing the execution requirements of the optimal control strategy.

[0072] The topology routing table of the cloud-edge collaborative architecture is retrieved to obtain the topology hop count from the cloud server to the target IoT device. This topology hop count includes not only the number of routers and switches traversed by the physical network layer, but also the hierarchical span of the logical control layer, i.e., the number of architectural layers in the "cloud-edge gateway-terminal device" structure. Based on the topology hop count, combined with the average single-hop transmission delay and packet loss rate of the current network link, the end-to-end transmission cost and command failure risk caused by the issuance of control commands along different paths are estimated. The larger the topology hop count, the more architectural layers the command penetrates, and the higher the uncertainty caused by transmission cost and link congestion, thus posing a substantial obstacle to the immediate effectiveness of high-frequency and high-critical control actions.

[0073] The frequency and criticality features of the control actions, as well as the topology hop count of the cloud-edge collaborative space, are fused into a multi-dimensional feature vector, which is then input into a preset hierarchical classifier. This hierarchical classifier adopts a nonlinear mapping model based on decision boundaries, and internally defines decision boundary partitioning rules corresponding to the collaborative control levels. In the feature space, the classifier evaluates the real-time matching degree of the issued instructions based on the joint constraints of frequency and topology hop count, and evaluates the reliability matching degree of the execution results based on the joint constraints of criticality and topology hop count. Through a nonlinear activation function, the confidence score of the comprehensive feature vector belonging to each collaborative control level is calculated, thus discretizing and mapping the continuous policy execution requirement features to a discrete hierarchical label space.

[0074] Based on the confidence scores output by the hierarchical classifier, a maximum pooling operation is performed to select the level with the highest confidence score as the collaborative control level that matches the current optimal control strategy. The collaborative control level includes cloud-based global overall control, edge gateway local control, and terminal device local self-determination control. If the frequency is extremely high and the number of topology hops is large, the classifier tends to output terminal device local self-determination control to avoid instruction lag caused by long-distance transmission. If the criticality is extremely high and involves cross-device spatial decoupling, the classifier tends to output cloud-based global overall control to utilize the abundant computing power and global perspective of the cloud to ensure the accuracy of core intervention. For local collaborative scenarios with moderate frequency and criticality and only one hop in the topology, edge gateway local control is output to achieve a trade-off between latency and computing power, thereby ultimately determining the execution subject and distribution path of the optimal control strategy.

[0075] Specifically, assuming that for the collision avoidance control of AGV_01, the cloud solution yields "Optimal control strategy C: fine-tuning the steering wheel, issuing a correction command every 0.1 seconds for 3 seconds" and for AGV_02, "Optimal control strategy D: emergency braking, instantaneously outputting the maximum braking current, a single action." During the feature extraction stage, the action interval of optimal control strategy C is extremely short, resulting in a "high-frequency" feature, but its action amplitude is small, and its criticality is low. Optimal control strategy D involves only one action, with an extremely low frequency, but it reaches the physical limits of the actuator, resulting in a "highly critical" feature.

[0076] A network topology routing table query in the cloud revealed that AGV_01 is currently within the coverage area of ​​a 5G macro base station, requiring a 3-hop journey to reach the cloud (terminal-base station-core network-cloud), resulting in a large number of hops and a transmission latency of approximately 50ms. AGV_02, on the other hand, is within the Wi-Fi coverage area of ​​the workshop edge gateway, requiring only one hop to reach the edge gateway and two hops to reach the cloud. Considering link quality, the cloud assessment concluded that: if high-frequency commands under Strategy A require three hops to reach the cloud and then be transmitted back, the 100ms round-trip latency would cause severe delays in fine-tuning commands, leading to AGV_01 exhibiting erratic oscillations; while Strategy B requires urgent execution, it only takes 5ms to reach the edge gateway with one hop and 30ms with two hops to reach the cloud.

[0077] The cloud inputs the feature vectors of the optimal control strategy C ([extremely high frequency, low criticality, 3 hops]) and the feature vectors of the optimal control strategy D ([extremely low frequency, extremely critical, 2 hops]) into the hierarchical classifier. The classifier's decision boundary rules begin to operate: For the optimal control strategy C, the classifier recognizes that the combination of "extremely high frequency" and "3 hops" severely violates the real-time requirements because network latency cannot support high-frequency fine-tuning, and low criticality means that global computing power in the cloud is not required. Therefore, it obtains the highest confidence on the decision boundary of "terminal self-determination". For the optimal control strategy D, although the 2-hop latency is acceptable, the "extremely critical" feature triggers the safety redundancy rule. If an emergency shutdown is handled by the edge or terminal with weaker computing power, a serious collision will occur if the computing crashes. Therefore, after comprehensive evaluation, the classifier leans more towards the more reliable cloud-based overall planning between "global cloud planning" and "local edge planning".

[0078] Based on the maximum confidence level output by the classifier, the cloud ultimately determines and outputs the collaborative control level: For the fine-tuning obstacle avoidance strategy of AGV_01, the determined and matched collaborative control level is the local autonomous control of the terminal device. The cloud will no longer issue high-frequency commands one by one, but will issue a meta-command to "activate the local high-frequency fine-tuning PID algorithm", which will be executed in closed loop by the microcontroller of AGV_01 using onboard sensors; For the emergency locking strategy of AGV_02, the determined and matched collaborative control level is the global coordination control of the cloud. This is because the cloud needs to coordinate and schedule the surrounding AGVs from AGV_03 to AGV_10 to bypass the obstacle while executing the locking, in order to prevent subsequent chain collisions. This strategy is directly issued by the cloud and synchronized with multiple vehicles to ensure the absolute global coordination of core safety actions.

[0079] refer to Figure 5 In step S14, the specific steps are as follows: S141: The cloud server distributes the optimal control policy to the matching collaborative control node through the message queue telemetry transmission protocol according to the determined collaborative control level, and uses timestamps and digital signatures to perform integrity verification and non-repudiation authentication on the policy messages; after the policy is executed, the observation data of the collaborative control node is collected in real time in an event-driven manner. S142: Synchronously acquire the current load of the cloud server and the current status of the IoT devices, calculate the actual control offset correction amount after the strategy is executed based on the observation data, dynamically adjust the control granularity and control command issuance cycle according to the current load of the cloud server and the current status of the IoT devices, and generate dynamic control content for the IoT devices in the next time slot by combining the fusion mechanism.

[0080] In the embodiments of this application, the cloud server distributes the optimal control policy to the matching collaborative control node through a message queue telemetry transmission protocol according to the determined collaborative control level, and uses timestamps and digital signatures to perform integrity verification and non-repudiation authentication on the policy messages; after the policy is executed, the observation data of the collaborative control node is collected in real time in an event-driven manner.

[0081] At this point, based on the collaborative control hierarchy determined in the previous steps—namely, global overall control in the cloud, local control at the edge gateway, or local self-determination control at the terminal device—the cloud server dynamically addresses the matching collaborative control node in the message distribution bus. It then uses the Message Queuing Telemetry Transport Protocol (MQTT) to execute the optimal control strategy distribution process. According to the determined hierarchy, the cloud server publishes a message encapsulating the optimal control strategy to a specific MQTT topic pre-agreed with the collaborative control node: for the cloud-overall hierarchy, it publishes to the global control topic; for the edge-local hierarchy, it publishes to the gateway proxy topic; and for the terminal-local hierarchy, it publishes to the device direct connection topic. Utilizing the MQTT protocol's QoS level mechanism, QoS 2 is configured for policy messages involving high-critical security to prevent executor jitter caused by repeated execution, and QoS 1 is configured for high-frequency fine-tuning strategies to balance real-time performance and delivery reliability. This achieves accurate routing and low-latency delivery of control commands in the cloud-edge-device architecture.

[0082] Before and after the optimal control policy message is sent, a two-way security verification mechanism based on timestamps and digital signatures is executed. Before being sent from the cloud, the cloud server uses its private key to digitally sign the policy message payload and the current high-precision timestamp, and encapsulates the signature value and timestamp into the message extension header. When the collaborative control node receives the MQTT message, it extracts the timestamp from the message and compares it with the node's local system clock. If the time difference exceeds a preset time-to-live threshold, the message is determined to be expired or a replay attack message and is discarded to mitigate the risk of delayed instruction execution caused by network latency. If the timestamp verification passes, the node further uses a preset cloud public key to decrypt and verify the digital signature, checking the consistency of the message payload hash value. Only when the hash consistency verification passes, confirming that the message has not been tampered with during transmission and that the signature source is indeed an authorized cloud server, is the legitimacy determined, realizing the integrity verification and non-repudiation authentication of the instruction, ensuring that control actions are executed within a secure and reliable closed loop.

[0083] After the collaborative control node authenticates the policy message, it parses it into control electrical signals or local logic rules that can be recognized by the underlying actuator and executes them. Simultaneously, an event-driven observation data acquisition mechanism is initiated, replacing the traditional high-frequency polling acquisition mode. At this time, event listeners are deployed on the sensing side of the IoT device, and dynamic trigger thresholds are set according to the expected response characteristics of the control policy. When the device state changes abruptly due to policy execution, or when the "state change rate, such as velocity derivative or position offset" exceeds the dynamic trigger threshold, the event listener immediately triggers data sampling and reporting. This mechanism effectively filters out redundant telemetry data under steady-state operation and concentrates network uplink bandwidth resources for transmitting key state transition observation data reflecting the control effect.

[0084] The collaborative control node will trigger the collection of observation data, including multimodal sensing parameters and network status indicators after execution, and then transmit it back to the cloud server's streaming computing engine via the MQTT protocol observation topic. After receiving the high-frequency burst event stream, the cloud streaming engine will perform streaming alignment and reconstruction based on the device identifier and timestamp to form an observation state sequence that represents the real physical feedback after the strategy is executed. This observation state sequence serves as a feedback signal for closed-loop control and is connected in real time to the evaluation module of the cloud collaborative space to compare the residual between the state pre-simulation trajectory and the actual observation trajectory.

[0085] Specifically, for the pre-determined collaborative control hierarchy, the cloud executes MQTT routing. For AGV_01, which is determined to be under "terminal local self-determination control," high-frequency fine-tuning is required. Instead of issuing a series of specific steering angle values, the cloud publishes the optimal strategy message "activate local high-frequency fine-tuning obstacle avoidance algorithm" to the MQTT topic factory / agv_01 / local_command, configuring QoS 1 to ensure that the meta-command quickly reaches the on-board MCU of AGV_01. For AGV_02 and surrounding AGVs, which are determined to be under "cloud-based global overall control," emergency braking and global replanning are required. The cloud publishes a strategy message containing specific braking intensity and replanning route to the global control topic factory / cluster / emergency_brake, configuring QoS 2 to ensure that even if the network fluctuates momentarily, the braking command will never be repeatedly sent, causing repeated wheel braking and unlocking oscillations.

[0086] Taking AGV_02 receiving an emergency brake command as an example; when the edge gateway forwards the MQTT message to AGV_02, AGV_02 reads the timestamp in the message header and finds that the timestamp differs from its own vehicle RTC clock by only 15ms, which is within the allowed 100ms survival time threshold. This eliminates the possibility that this is an "old command" that was delayed due to network congestion a few seconds ago. Executing an old brake command would cause the AGV to stop abruptly at the wrong position. AGV_02 calls the preset cloud public key to decrypt the digital signature and calculates that the message hash value is completely consistent with the received hash value. This confirms that the brake command was indeed issued by the cloud server and has not been tampered with by malicious interference nodes in the factory. Only after completing the non-repudiation authentication does the drive motor controller execute the braking action.

[0087] After AGV_01 executes the "local high-frequency fine-tuning" strategy, if the traditional method of polling and reporting coordinates every 10ms is used, the massive amount of high-frequency telemetry data will instantly exceed the throughput capacity of the wireless channel, causing severe network congestion and transmission queue overflow. At this time, the event-driven mechanism comes into play: AGV_01 locally monitors its own heading angle deviation and lateral offset, and only when the lateral offset changes by more than 2 cm or the heading angle changes by more than 1 degree is it considered a "valid event". When AGV_01 fine-tunes the steering wheel and produces a displacement of more than 2 cm, the listener immediately triggers a status report; while during the period when AGV_01 is smoothly straight-line fine-tuning, it remains silent and does not send data.

[0088] After executing the emergency braking strategy coordinated by the cloud, AGV_02's onboard sensors detected a slight sideslip caused by the slippery road surface. This sideslip caused a sudden change in the heading angle, triggering an event-driven report. After receiving this sudden sideslip data through the MQTT observation topic, the cloud-based Flink streaming engine immediately compared it with the previously rehearsed "smooth deceleration to zero" trajectory in the cloud collaborative space. It found a significant residual between the two, which served as a closed-loop feedback signal and was connected back to the cloud's evaluation module in real time. This indicated that the current braking strategy in the cloud had caused unexpected instability on the slippery road surface, and it was necessary to immediately proceed to the next step S142. Based on the cloud load and the instability state, the control content was dynamically adjusted, such as adding anti-lock braking commands.

[0089] Furthermore, the system synchronously acquires the current load of the cloud server and the current status of IoT devices, calculates the actual control offset correction amount after strategy execution based on the observation data, and dynamically adjusts the control granularity and control command issuance cycle according to the current load of the cloud server and the current status of IoT devices. Combined with the fusion mechanism, it generates dynamic control content for IoT devices in the next time slot, further controlling the collaborative control level. It fully considers the observation data, the current load of the cloud server, and the current status of IoT devices, thus improving the accuracy of the dynamic control content of IoT devices.

[0090] At this time, during the operation of the cloud server, the current load status of the cloud server itself is synchronously obtained through the resource monitoring daemon process. The current load status includes the CPU utilization, memory usage, and message backlog depth of the streaming computing engine. Simultaneously, the current state of the IoT device is extracted through the observation data collected in real time in the previous steps. The current state includes the physical wear and tear of the device actuator, the signal-to-noise ratio of the sensor, and the instantaneous jitter rate of the network link. Further, based on the observation data, the offset calculation model of the previous step S122 is re-substituted to calculate the actual control offset correction amount of the system after the strategy is executed. This actual control offset correction amount characterizes the degree of safety deviation remaining in the system or the degree of oscillation deviation caused by strategy overshoot after the optimal control strategy is executed under real physical and environmental disturbances, providing a quantitative basis for the subsequent dynamic adaptive adjustment of control parameters.

[0091] Based on the current load of the cloud server and the current state of the IoT devices, and combined with the actual control offset correction, the control granularity is dynamically adjusted. Here, control granularity is defined as the step size of the device state change driven by a single control command and the spatial decoupling dimension involved. When the current load of the cloud server exceeds the high load threshold and the actual control offset correction is less than the safe convergence threshold, the system determines that the current macroscopic situation has stabilized. At this point, a control granularity coarsening strategy is executed, increasing the adjustment step size of the control action and reducing the multi-dimensionally coupled "fine-grained control commands, such as independently controlling the difference in speed between the left and right wheels," into "coarse-grained control commands, such as the target speed and heading angle of the entire vehicle," to reduce the computational complexity and command generation overhead in the cloud. Conversely, when the cloud load margin is sufficient and the actual control offset correction indicates that the system is still in a high-frequency oscillation or incompletely converged unstable range, a control granularity refinement strategy is executed, reducing the adjustment step size and decoupling multi-dimensional control variables, outputting a refined control law to accelerate system convergence.

[0092] The control command issuance cycle is dynamically adjusted based on the current load of the cloud server and the current state of the IoT device. The issuance cycle is the logical time interval between two consecutive control strategy issuances. When the current state of the IoT device indicates an increase in the real-time jitter rate of its network link or an increase in the message backlog depth of the cloud server, maintaining a high frequency and short cycle issuance will cause commands to be queued out of order in the transmission queue or cause command delays and overlapping execution due to network congestion. At this time, a cycle lengthening mechanism is triggered to proportionally increase the issuance cycle, trading time for space to ensure that each command is reliably transmitted and correctly executed in the link. Conversely, when the link is smooth and the cloud computing power is sufficient, and the actual control offset correction reflects that the device is experiencing transient changes, a cycle compression mechanism is triggered to shorten the issuance cycle to the millisecond level, using high-frequency strategy iteration to combat rapid external disturbances.

[0093] The adjusted dynamic control granularity and dynamic control command issuance cycle are used as constraint parameters. Combined with the actual control offset correction amount and state pre-simulation trajectory of the current time slot, a fusion mechanism is introduced to generate dynamic control content for IoT devices in the next time slot. The fusion mechanism is as follows: using the actual control offset correction amount as the correction benchmark, within the action space defined by the dynamic control granularity, a numerical optimization algorithm or feedforward compensation network is used to solve for the control action sequence that can maximize the offset under the current dynamic constraints. Based on the dynamic control command issuance cycle, the sequence is resampled in the time dimension and smoothed with a zero-order hold to eliminate the command step jump caused by the cycle switching. The control command that has been corrected and smoothed by time resampling is encapsulated into dynamic control content for the next time slot, realizing elastic adaptive adjustment based on real-time situation under the closed-loop control architecture.

[0094] Specifically, after AGV_02 executed the emergency lock-up strategy issued by the cloud, the cloud resource monitoring daemon found that the current CPU utilization had reached 85%, the message backlog depth had reached 2000 messages, and the cloud was currently under high load. At the same time, the observation data reported by AGV_02 showed that its wheels had slightly skidded due to the oil on the road surface, and the Wi-Fi signal strength was weak, indicating that the equipment was in poor condition. Based on these observation data, the cloud recalculated the actual control offset correction and found that although the emergency lock-up command was executed, due to the skidding, the actual trajectory of AGV_02 deviated from the safety envelope by 0.15 meters, and the vehicle's heading had a 3-degree yaw angle. These 0.15 meters and 3 degrees are the residual offset correction after the strategy was executed, indicating that the system has not yet safely converged.

[0095] Due to the extremely high load on the cloud platform and the fact that the 0.15-meter residual offset of AGV_02 is still within a controllable range and less than the dangerous threshold of 0.2 meters, the cloud platform has decided to implement a coarsening control granularity strategy. Originally, in order to correct the 3-degree yaw angle, the cloud platform needed to calculate the fine differential torque of the left and right wheels of AGV_02 separately, which required a lot of computing power. Now, the cloud platform will coarsen the control granularity and directly generate a simplified steering correction command of "left wheel speed limit 0.1m / s, right wheel speed limit 0.15m / s", without performing torque distribution calculations.

[0096] In response to the combined issues of weak network signal and high jitter rate in AGV_02, as well as message backlog in the cloud, the cloud dynamically adjusts the control command sending cycle. Originally, the strategy for side-slip correction had a sending cycle of 100ms (high frequency). However, under the current conditions, within 100ms, the previous command might not have reached AGV_02 before the next command is sent, causing command backlog in the router or cumulative fast-forwarding during execution. Therefore, the cloud triggers a cycle lengthening mechanism, adaptively adjusting the sending cycle to 500ms. This allows AGV_02 ample time to receive, verify, and execute a correction action.

[0097] The cloud-based system uses a coarse control granularity of "differentiated speed limits for left and right wheels" and a lengthened "500ms" transmission cycle as constraints, integrating a 0.15m and 3-degree offset correction to generate dynamic control content for the next time slot. A feedforward compensation network calculates the left and right wheel speed difference that can offset the 3-degree yaw angle within 500ms, and a zero-order hold smooths the command to prevent mechanical shock caused by the command abruptly jumping from 0 to 0.15m / s during cycle switching. The cloud generates and transmits dynamic control content containing "left wheel 0.1m / s, right wheel 0.15m / s, continuous action for 500ms," which corrects sideslip and perfectly adapts to the current high load on the cloud and the weak network state of the AGV.

[0098] Please see Figure 6 The dynamic control system for cloud-based collaborative IoT devices is applied to the aforementioned dynamic control method for cloud-based collaborative IoT devices; the dynamic control system for cloud-based collaborative IoT devices includes: The dynamic feature module 21 is used to receive multi-source heterogeneous stream data from IoT devices in the cloud server, and construct the global state space of IoT devices by combining it with a preset spatiotemporal correlation graph. It performs feature aggregation and dimensionality reduction on the global state space and extracts dynamic feature vectors containing temporal dependency characteristics and spatial coupling characteristics. The cloud collaboration module 22 is used to input the dynamic feature vector into the corresponding cloud collaboration space, generate the state prediction trajectory of the IoT device in the future time domain during the feature evolution process, and determine the corresponding control offset by comparing it with the preset safe operation trajectory. The collaborative control module 23 is used to solve the optimal control strategy with the joint optimization objectives of delay minimization and control stability rate based on the dynamic feature vector and state evolution trajectory when the control offset meets the dynamic collaborative triggering conditions, and to determine the collaborative control level that matches the optimal control strategy. The dynamic control module 24 is used by the cloud server to distribute the optimal control strategy to the corresponding collaborative control node according to the collaborative control hierarchy, and to collect observation data after the strategy is executed in real time. It further combines the current load of the cloud server and the current status of the IoT device to determine the dynamic control content of the IoT device.

[0099] It should be noted that although multiple modules are mentioned in the detailed description above, this division is not mandatory; in fact, according to the embodiments of this disclosure, the features and functions of two or more modules or described above can be embodied in one module; conversely, the features and functions of one module described above can be further divided into multiple modules to be embodied.

[0100] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein; this application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein; the specification and embodiments are to be considered exemplary only.

Claims

1. A dynamic control method of an Internet of Things device based on cloud collaboration, characterized in that, include: In the cloud server, multi-source heterogeneous stream data from IoT devices are received, and a global state space of IoT devices is constructed by combining it with a preset spatiotemporal correlation graph. Feature aggregation and dimensionality reduction are performed on the global state space to extract dynamic feature vectors containing temporal dependency characteristics and spatial coupling characteristics. The dynamic feature vector is input into the corresponding cloud collaborative space, and the state prediction trajectory of the IoT device in the future time domain is generated during the feature evolution process. The corresponding control offset is determined by comparing it with the preset safe operation trajectory. When the control offset meets the dynamic collaborative triggering condition, the cloud server solves the optimal control strategy with the joint optimization objectives of delay minimization and control stability rate based on the dynamic feature vector and state evolution trajectory, and determines the collaborative control level that matches the optimal control strategy. The cloud server distributes the optimal control strategy to the corresponding collaborative control node according to the collaborative control hierarchy, and collects observation data after the strategy is executed in real time. It further combines the current load of the cloud server and the current status of the IoT device to determine the dynamic control content of the IoT device.

2. The cloud collaboration based dynamic control method of IoT devices as claimed in claim 1 wherein, In the cloud server, multi-source heterogeneous stream data from IoT devices is received, and a global state space for the IoT devices is constructed by combining it with a preset spatiotemporal correlation graph. Feature aggregation and dimensionality reduction are performed on the global state space to extract dynamic feature vectors containing temporal dependency and spatial coupling characteristics, including: During the operation of IoT devices, the cloud server receives multi-source heterogeneous stream data reported by IoT devices in real time based on the streaming computing engine. This multi-source heterogeneous stream data includes multimodal sensing parameters and network status indicators. Combined with graph attention network, a preset spatiotemporal correlation graph is constructed with devices as nodes and physical topology and logical communication links between devices as edges. The multi-source heterogeneous stream data is mapped to the spatiotemporal correlation graph, thereby constructing the global state space of IoT devices. 3.The cloud collaboration based dynamic control method of IoT devices according to claim 2, wherein, The process of receiving multi-source heterogeneous stream data from IoT devices in a cloud server, constructing a global state space for IoT devices by combining it with a preset spatiotemporal correlation graph, performing feature aggregation and dimensionality reduction on the global state space, and extracting dynamic feature vectors containing temporal dependency and spatial coupling characteristics, further includes: An autoencoder with a spatiotemporal masking mechanism is obtained to perform feature aggregation and dimensionality reduction on the global state space. During the aggregation process, the spatial coupling characteristics of the IoT device cluster are captured by an adaptive graph pooling layer, and the temporal dependency characteristics of node state evolution are extracted by a temporal convolutional network. In this way, a dynamic feature vector containing temporal dependency characteristics and spatial coupling characteristics is extracted.

4. The dynamic control method for IoT devices based on cloud collaboration according to claim 1, characterized in that, The process of inputting dynamic feature vectors into the corresponding cloud-based collaborative space, generating a state prediction trajectory of IoT devices in the future time domain during feature evolution, and determining the corresponding control offset by comparing it with a preset safe operation trajectory includes: The cloud collaboration space deployed in the cloud is obtained. Dynamic feature vectors are input into the cloud collaboration space. Based on the fusion architecture of long short-term memory network and graph neural network, state evolution events are determined. In the state evolution events, the nonlinear state transition of dynamic feature vectors is deduced to generate multi-step state prediction trajectories of IoT devices in the future time domain.

5. The dynamic control method for IoT devices based on cloud collaboration according to claim 4, characterized in that, The step of inputting dynamic feature vectors into the corresponding cloud-based collaborative space, generating a state prediction trajectory of IoT devices in the future time domain during feature evolution, and determining the corresponding control offset by comparing it with a preset safe operating trajectory, further includes: A preset safe operating trajectory is obtained, which is generated by the physical boundary and operating condition constraints of the IoT device. The state pre-simulation trajectory and the safe operating trajectory are dynamically aligned and compared in the same space. The weighted sum of the distance and trajectory slope deviation between the state pre-simulation trajectory and the safe operating trajectory in the same space is determined, and the weighted sum is quantified as the control offset degree characterizing the degree of deviation of the system from the safe state.

6. The dynamic control method for IoT devices based on cloud collaboration according to claim 1, characterized in that, When the control offset meets the dynamic collaborative triggering condition, the cloud server, based on the dynamic feature vector and state evolution trajectory, solves for the optimal control strategy with the joint optimization objectives of delay minimization and control stability, and determines the collaborative control hierarchy matching the optimal control strategy, including: The control offset is monitored in real time. When the control offset exceeds the preset offset threshold, it is determined that the dynamic collaborative triggering condition is met. The cloud server constructs a Markov decision process based on dynamic feature vectors and state evolution trajectories. The solution of the control strategy is transformed into a multi-objective optimization problem with the joint optimization objectives of minimizing end-to-end communication delay and maximizing the dynamic response control stability rate of the system. Further, the solution is explored by combining reinforcement learning mechanism, and the non-dominated solution set is selected to determine the optimal control strategy.

7. The dynamic control method for IoT devices based on cloud collaboration according to claim 6, characterized in that, When the control offset meets the dynamic collaborative triggering condition, the cloud server, based on the dynamic feature vector and state evolution trajectory, solves for the optimal control strategy with the joint optimization objectives of delay minimization and control stability, and determines the collaborative control level matching the optimal control strategy, further including: The frequency and criticality features of control actions in the optimal control strategy are extracted. Combined with the topological hop count of the cloud-edge collaborative space, a hierarchical classifier is used to determine the collaborative control level that matches the optimal control strategy. The collaborative control level includes global overall control in the cloud, local control in the edge gateway, and local self-determination control in the terminal device.

8. The dynamic control method for IoT devices based on cloud collaboration according to claim 1, characterized in that, The cloud server distributes the optimal control strategy to the corresponding collaborative control nodes according to the collaborative control hierarchy, and collects observation data after the strategy execution in real time. Furthermore, it combines the current load of the cloud server and the current state of the IoT devices to determine the dynamic control content of the IoT devices, including: The cloud server distributes the optimal control policy to the matching collaborative control nodes through a message queue telemetry transmission protocol according to the determined collaborative control level, and uses timestamps and digital signatures to perform integrity verification and non-repudiation authentication on the policy messages; after the policy is executed, the observation data of the collaborative control nodes are collected in real time in an event-driven manner.

9. The dynamic control method for IoT devices based on cloud collaboration according to claim 8, characterized in that, The cloud server distributes the optimal control strategy to the corresponding collaborative control nodes according to the collaborative control hierarchy, and collects observation data after the strategy execution in real time. It further combines the current load of the cloud server and the current state of the IoT devices to determine the dynamic control content of the IoT devices, and also includes: The system synchronously acquires the current load of the cloud server and the current status of IoT devices, calculates the actual control offset correction amount after the strategy is executed based on the observation data, dynamically adjusts the control granularity and control command issuance cycle according to the current load of the cloud server and the current status of IoT devices, and generates dynamic control content for IoT devices for the next time slot by combining the fusion mechanism.

10. A dynamic control system for IoT devices based on cloud collaboration, characterized in that, The dynamic control system for cloud-based collaborative IoT devices is applied to the dynamic control method for cloud-based collaborative IoT devices as described in any one of claims 1-9; The dynamic control system for the cloud-based collaborative IoT device includes: The dynamic feature module is used to receive multi-source heterogeneous stream data from IoT devices in the cloud server, and construct the global state space of IoT devices by combining it with a preset spatiotemporal correlation graph. It then performs feature aggregation and dimensionality reduction on the global state space to extract dynamic feature vectors containing temporal dependency characteristics and spatial coupling characteristics. The cloud collaboration module is used to input dynamic feature vectors into the corresponding cloud collaboration space, generate the state prediction trajectory of IoT devices in the future time domain during the feature evolution process, and determine the corresponding control offset by comparing with the preset safe operation trajectory. The collaborative control module is used to solve the optimal control strategy with the joint optimization objectives of delay minimization and control stability rate based on dynamic feature vectors and state evolution trajectories when the control offset meets the dynamic collaborative triggering conditions, and to determine the collaborative control level that matches the optimal control strategy. The dynamic control module is used by the cloud server to distribute the optimal control strategy to the corresponding collaborative control node according to the collaborative control hierarchy, and to collect observation data after the strategy is executed in real time. It further combines the current load of the cloud server and the current status of the IoT device to determine the dynamic control content of the IoT device.