Adaptive hierarchical scheduling method fusing TSN and federated learning system

By introducing an adaptive hierarchical scheduling method into the TSN system, optimizing client selection and data stream injection, the network congestion and latency problems under large-scale heterogeneous devices are solved, and fast convergence and efficient training are achieved.

CN120980040APending Publication Date: 2025-11-18CHONGQING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511098921.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing TSN systems suffer from insufficient scheduling resources, network congestion, and task delays in large-scale, heterogeneous device scenarios. Existing methods lack coordinated optimization of communication and training processes, resulting in limited overall system efficiency.

Method used

An adaptive hierarchical scheduling method integrating TSN and federated learning systems is adopted. The scheduling strategy is generated by the hierarchical scheduler. Combined with spatiotemporal state encoder and reinforcement learning, the client selection and data stream injection are optimized to achieve dynamic network resource management.

Benefits of technology

It improves the training rate of federated learning systems, enabling them to quickly converge to the specified accuracy in industrial TSN scenarios, mitigating the risk of queue overflow and improving system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980040A_ABST
    Figure CN120980040A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive hierarchical scheduling method fusing a TSN and a federated learning system, and belongs to the field of industrial Internet, and the scheduling method comprises the following specific steps: I, constructing a time-sensitive industrial Internet of Things framework, and initializing various parameters of the time-sensitive industrial Internet of Things framework; iI, the hierarchical scheduler generates a latest scheduling strategy, and when each round of training begins, based on the scheduling strategy, participated clients of the current round of training are selected from the clients and the TSN transmission mode of the participated clients is determined; the method has the capabilities of sensing the computing capability isomerism, predicting the TSN link load and relieving the queue overflow risk, effectively improves the training rate of the federated learning system, and can realize the rapid convergence of the federated learning system under the specified precision in the industrial TSN scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of industrial internet, and in particular to an adaptive hierarchical scheduling method for a system combining TSN and federated learning. BACKGROUND

[0002] With the continuous promotion of industrial internet in manufacturing, energy, transportation and other key industries, it is a core requirement to ensure the real-time, reliability and deterministic transmission of network systems. As a key technology based on IEEE standards, Time Sensitive Networking (TSN) plays a supporting role in time-sensitive tasks in industrial data transmission by introducing precise clock synchronization and traffic scheduling protocols. TSN can effectively reduce communication delay and jitter, meeting the stringent requirements of end-to-end transmission delay and determinism in industrial automation.

[0003] To further improve the coordination efficiency of industrial intelligent systems, researchers have proposed in recent years to combine TSN and federated learning mechanisms to build a system architecture with learning ability and transmission scheduling optimization ability. In particular, in scenarios where large-scale distributed devices are deployed on the edge side, each terminal device not only has computing power, but also needs to upload model updates and other data through the TSN network, thereby realizing an intelligent learning closed loop of edge-cloud collaboration. However, existing TSN systems still have problems such as insufficient scheduling resources, network congestion and task delay when facing large-scale, heterogeneous devices uploading learning streams. For example, if multiple devices concurrently transmit data in the same time window, it may cause the queue of critical tasks to overflow or even information loss, thereby affecting the overall efficiency and convergence speed of model training. In addition, due to the heterogeneity of end device computing resources, uneven data quality, and different path loads, a unified static scheduling strategy cannot adapt to the dynamically changing network state and learning task requirements.

[0004] To address these challenges, several studies or patents have proposed optimization schemes with different focuses. However, most existing methods lack a collaborative optimization mechanism between the communication and training processes, and fail to integrate the scheduling strategies of data transmission and model training from a global perspective, resulting in limited overall system efficiency. For example, the client sampling algorithm based on reinforcement learning proposed by multiple federated learning patents and researches (such as FedCS, Oort) emphasizes selecting participating nodes based on indicators such as computing power, data quality, and network delay to speed up training convergence. However, it ignores the dynamic and congestion effects of network resources in the communication process, especially in TSN, where multiple "high-quality" nodes being scheduled simultaneously may cause serious transmission conflicts and queue overflow, leading to training interruption or failure.

[0005] CN113630893A discloses a 5G and TSN joint scheduling method based on wireless channel information, which mainly adjusts the retransmission factor, optimizes the priority mapping and physical layer transmission order from the perspective of physical layer and link layer to improve the determinism and throughput of TSN transmission. However, it does not consider the dynamic demand of upper layer tasks (such as distributed learning tasks) for network resources, and the device behavior is passive, so the scheduling decision does not have perceptivity and adaptivity. Therefore, we propose an adaptive hierarchical scheduling method for fusion of TSN and federated learning system. SUMMARY

[0006] The purpose of the present application is to solve the defects in the prior art, and an adaptive hierarchical scheduling method for fusion of TSN and federated learning system is proposed.

[0007] In order to achieve the above purpose, the present application adopts the following technical scheme:

[0008] The adaptive hierarchical scheduling method for fusion of TSN and federated learning system has the following specific steps:

[0009] I: Construct a time-sensitive industrial Internet of Things framework and initialize the parameters of the time-sensitive industrial Internet of Things framework;

[0010] II: The hierarchical scheduler generates the latest scheduling strategy, and at the beginning of each round of training, based on the scheduling strategy, selects the participating client in this round of training from the client and determines its TSN transmission mode;

[0011] III: The central server broadcasts the current global model parameters to each participating client, and each participating client performs local training according to the received parameter information;

[0012] IV: Encapsulate the local parameters obtained in each round of training into FL flow and upload, and update the global model parameters in the central server based on each FL flow;

[0013] V: Record the training time of each round, collect the results of the hierarchical scheduler, and optimize the hierarchical scheduler decision using the experience replay mechanism;

[0014] VI: Repeat the selection of client, local training, global model parameter update and strategy update until the global model reaches the specified precision or reaches the preset maximum iteration number.

[0015] As a further scheme of the present application, the time-sensitive industrial Internet of Things framework of step I specifically includes a device layer, a hierarchical scheduler, a TSN network, a TSN CQF circular queue and a field layer.

[0016] The device layer comprises a plurality of edge devices in an industrial Internet of Things, i.e., clients, for collecting data and local computing, and each selected client in each round completes local computing and encapsulates data into a FL flow, and according to the decision result of the hierarchical scheduler, the FL flow is transmitted to the TSN network at different time points;

[0017] The hierarchical scheduler is composed of a space-time state encoder, an upper-layer client sampling strategy and a lower-layer injection control strategy, and is respectively responsible for global state modeling, joint optimization of client selection and injection time slot allocation;

[0018] The space-time state encoder is used to extract information including the computing capability of the federal learning client, the data quality and the TSN path congestion, and generate a global state representation; the upper-layer client sampling strategy is used for the federal learning user selection strategy, and the client with stronger computing capability and data quality meeting the preset condition is selected to participate in the training in this round; and the lower-layer injection control strategy is used to optimize the real-time flow injection strategy of the edge device, i.e., the time slot offset of the injection TSN network, so as to perform real-time transmission of the TSN network.

[0019] The TSN network comprises a TSN gateway and a TSN switch, which are used to receive and forward the FL flow injected by the edge device, and perform delay deterministic forwarding on the FL flow;

[0020] The TSN gateway is used to receive the FL flow and transmit it to the central server; the TSN switch is used to perform multi-hop forwarding on the FL flow injected into the TSN network; and the TSN gateway and the TSN switch utilize the TSN CQF circular queue to realize deterministic transmission of the data packet.

[0021] The TSN CQF circular queue is composed of a sending queue and a receiving queue;

[0022] The field layer is composed of an edge server, a controller and a traditional field-level device, wherein the edge server serves as a central server of the federal learning to aggregate parameters, and the controller makes hierarchical decisions on user selection and transmission delay.

[0023] As a further scheme of the application, a forwarding time slot aligned with the transmission time slot is set, and the transmission time slot and the forwarding time slot are collectively referred to as a time slot; in a time slot, the receiving queue in the TSN CQF circular queue only receives the data packet to be forwarded, and the sending queue only sends the data packet in the queue; in the next time slot, the two queues are cyclically switched, i.e., the original receiving queue is converted into the sending queue, and the data packet received in the previous time slot is sent out, while the sending queue in the previous time slot is used to receive the data packet in the current time slot.

[0024] As a further scheme of the present application, the parameters of the time-sensitive industrial internet of things framework in step I specifically include a client subset size, a client local computing frequency, a data volume size, model initial parameters, a time slot length, and a queue length.

[0025] As a further scheme of the present application, the specific steps for the layered scheduler to generate the latest scheduling strategy in step II are as follows:

[0026] S1.1: The layered scheduler converts the total training time of federated learning into a joint optimization problem of user sampling + TSN scheduling, models the joint optimization problem as a semi-Markov decision process, and then generates optimal strategies for the upper and lower layers through sampling layered reinforcement learning to achieve the optimization objective of shortening the total training data while reaching the specified precision, and establishes an optimization problem;

[0027] S1.2: Each edge device updates the client state through a space-time state encoder according to the current TSN network state and historical information before each round of training begins, and establishes the upper-layer client sampling strategy and the lower-layer injection control strategy of the layered scheduler based on the optimization problem, the reinforcement learning strategy, and the semi-Markov decision process framework.

[0028] As a further scheme of the present application, the specific form of the joint optimization problem of user sampling + TSN scheduling in S1.1 is as follows:

[0029]

[0030] In the formula, ζ and δ represent relevant constants, respectively; p and Δ represent the upper-layer client sampling strategy and the lower-layer injection control strategy, respectively; q and G represent the proportion of client data volume and the local gradient range, respectively; represents the training time of each round.

[0031] As a further scheme of the present application, the constraints of the joint optimization problem of user sampling + TSN scheduling include: for each queue e on each TSN switch, the queue length Q e,t (p, Δ) does not exceed the upper limit Q max , and the sum of all user sampling probabilities is one, and the time offset of each FL flow does not exceed the upper limit Δ max .

[0032] As a further scheme of the present application, the layered reinforcement learning method is adopted in the joint optimization problem of user sampling + TSN scheduling, and the entire Markov decision is decomposed as follows:

[0033]

[0034] Decomposition transforms the whole learning framework into a semi-Markov decision process with two different decision time scales, and since the ultimate goal is to minimize the overall time, the upper Q function is expressed as follows:

[0035]

[0036] where, γ h represents a discount factor, and and represent the start state and start action, respectively; C r represents an instantaneous cost expression, which is specifically expressed as follows:

[0037]

[0038] where the lower Q function has a similar expression as the upper Q function, but the instantaneous cost calculation formula is expressed as follows:

[0039]

[0040] where, C p represents a queue overflow penalty factor; and represents the transmission time of each FL flow after a given user offset.

[0041] As a further scheme of the present application, the client state specifically includes computing power, data quality, characteristics of data flow (such as data packet size, start time, transmission path, etc.), and congestion of the network.

[0042] As a further scheme of the present application, the specific steps of the local training of each participating client according to the received parameter information in step III are as follows:

[0043] S2.1: At the beginning of each training round, the hierarchical scheduler collects all client states, and then generates client sampling probabilities based on the upper client sampling strategy to obtain participating clients in this round. The central server downloads the current global model parameters to each participating client and replaces the original parameters of the local model of each participating client.

[0044] S2.2: Each participating client uses the local data collected by the sensor to iteratively train the local model, and then encapsulates the local model parameters after training into a time-sensitive FL flow and transmits the result to the lower injection control strategy.

[0045] S2.3: The lower injection control strategy receives information of each participating client and collects local observation data of each participating client, and then updates the lower injection control strategy state according to each local observation data and generates transmission decisions corresponding to each participating client.

[0046] As a further scheme of the present application, the specific steps of updating the global model parameters in the central server are as follows:

[0047] S3.1: Each participating client uploads the encapsulated FL flow respectively, the hierarchical scheduler executes the selected transmission decision, injects the FL flow into the TSN network, and observes the results, wherein the observation results include whether the injection is successful and the influence on the network state, and the transmission time of the flow is recorded;

[0048] S3.2: According to the injection result, the lower layer injection control strategy obtains the corresponding reward signal, which represents the good or bad of the action, stores the current state of each participating client, the executed action, the obtained reward and the subsequent state as experience tuples, and puts them into the lower layer replay buffer;

[0049] S3.3: When all participating clients in this round complete the transmission, the parameter aggregation is performed in the central server, then the total time consumption of the local training and the user with the slowest transmission time in this round is taken as the training time of this round, and the upper reward signal is obtained according to the transmission time of this round and the sampling probability obtained by the upper client sampling strategy, the current state of each participating client, the executed action, the obtained reward and the subsequent state are stored as experience tuples, and put into the upper layer replay buffer.

[0050] As a further scheme of the present application, the specific steps of optimizing the decision of the hierarchical scheduler in step V are as follows:

[0051] S4.1: Determine whether the accumulated experience in the upper layer replay buffer and the lower layer replay buffer reaches the preset threshold, if not, continue to interactively collect;

[0052] S4.2: If the accumulated experience in the upper layer replay buffer and the lower layer replay buffer reaches the preset threshold, training samples are extracted from the upper layer replay buffer and the lower layer replay buffer according to the priority sampling strategy respectively;

[0053] S4.3: According to the policy gradient theorem, update the upper client sampling strategy parameters, then train the lower layer injection control strategy of each participating client through non-cumulative Bellman update, and establish the corresponding action function with the goal of minimizing the maximum delay in a single round, and update the lower layer injection control strategy parameters based on the action function;

[0054] S4.4: When the iteration period reaches the preset update period, the target network parameters are updated to the current estimated network parameters.

[0055] As a further scheme of the present application, the specific calculation formula of updating the upper client sampling strategy parameters in S4.3 is as follows:

[0056]

[0057] In the formula, alpha h represents the upper layer learning rate; Q h (s r , p r ) represents the action value function of the upper layer client sampling strategy, which is estimated by using a deep double Q network and realizes stable learning by combining target network parameters, and the action value function is responsible for evaluating the action p under the state s;

[0058] The specific calculation formula of updating the lower layer injection control strategy parameter S4.3 is as follows:

[0059]

[0060] In the formula, represents the action value function of the lower layer injection control strategy.

[0061] Compared with the prior art, the beneficial effects of the present application are:

[0062] The present application generates the upper layer client sampling strategy and the lower layer TSN injection control strategy through a hierarchical reinforcement learning framework. Before each round of training starts, the system collects the computing power, data characteristics and network state of the device. After updating the client state, the participating training client is selected according to the upper layer strategy. After the local training of the client is completed, the injection timing of the FL flow is determined by the lower layer strategy and transmitted to the TSN network. After the training and transmission are completed, the system aggregates the parameters on the server and updates the upper and lower layer strategies based on the training time consumption and the strategy effect. The strategy performance is improved by using experience replay and priority sampling, and the target network is updated periodically. It has the ability to perceive the heterogeneity of computing power, predict the load of TSN link and relieve the risk of queue overflow. The training rate of the federated learning system is effectively improved, and the federated learning system under the industrial TSN scene can achieve fast convergence to reach the specified precision. BRIEF DESCRIPTION OF DRAWINGS

[0063] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, which together with the embodiments of the present application, serve to explain the present application, and do not constitute a limitation on the present application.

[0064] Fig. 1 The flow chart of the adaptive hierarchical scheduling method of the fusion TSN and federated learning system proposed by the present application;

[0065] Fig. 2 The global structure diagram of the adaptive hierarchical scheduling method of the fusion TSN and federated learning system proposed by the present application;

[0066] Fig. 3 The hierarchical scheduling mechanism structure diagram of the adaptive hierarchical scheduling method of the fusion TSN and federated learning system proposed by the present application. DETAILED DESCRIPTION

[0067] Referring to Figs. 1-3 , the adaptive hierarchical scheduling method fuses TSN with the federated learning system, and the specific steps are as follows:

[0068] A time-sensitive industrial Internet of Things framework is constructed, and parameters of the time-sensitive industrial Internet of Things framework are initialized.

[0069] It should be further explained that the time-sensitive industrial Internet of Things framework specifically includes a device layer, a hierarchical scheduler, a TSN network, a TSN CQF circular queue, and a field layer.

[0070] The device layer includes a plurality of edge devices in the industrial Internet of Things, i.e., clients, for collecting data and local computing. After the selected client completes the local computing in each round, the client encapsulates the transmission data as an FL flow and transmits the FL flow to the TSN network at different times according to the decision result of the hierarchical scheduler.

[0071] The hierarchical scheduler is composed of a space-time state encoder, an upper-layer client sampling strategy, and a lower-layer injection control strategy, which are respectively responsible for global state modeling, joint optimization of client selection and injection time slot allocation. The space-time state encoder is used to extract information including the computing capability of the federated learning client, the data quality, and the congestion of the TSN path, and generate a global state representation. The upper-layer client sampling strategy is used for the federated learning user selection strategy, which selects clients with stronger computing capability and data quality meeting the preset conditions to participate in the training in this round. The lower-layer injection control strategy is used to optimize the real-time flow injection strategy of the edge device, i.e., the time slot offset of the injection into the TSN network, for real-time transmission of the TSN network.

[0072] The TSN network includes a TSN gateway and a TSN switch, which are used to receive and forward the FL flow injected by the edge device, and perform delay deterministic forwarding. The TSN gateway is used to receive the FL flow and transmit it to the central server. The TSN switch is used to perform multi-hop forwarding of the FL flow injected into the TSN network. The TSN gateway and the TSN switch use the TSN CQF circular queue to realize deterministic transmission of data packets.

[0073] The TSN CQF circular queue is composed of a sending queue and a receiving queue, and a forwarding time slot aligned with the transmission time slot is set. The transmission time slot and the forwarding time slot are collectively referred to as a time slot. In a time slot, the receiving queue in the TSN CQF circular queue only receives data packets to be forwarded, and the sending queue only sends data packets in the queue. In the next time slot, the two queues are cyclically rotated, i.e., the original receiving queue is converted into a sending queue, and the data packets received in the previous time slot are sent out, while the sending queue in the previous time slot is used to receive data packets in the current time slot.

[0074] The field layer is composed of edge servers, controllers and traditional field-level devices, wherein the edge server serves as a central server of federated learning to aggregate parameters, and the controller makes hierarchical decisions on user selection and transmission delay.

[0075] In addition, it should be noted that the parameters of the time-sensitive industrial Internet of Things framework specifically include a client subset size, a client local computing frequency, a data size, model initial parameters, a time slot length and a queue length.

[0076] The hierarchical scheduler generates the latest scheduling strategy at the beginning of each round of training, and based on the scheduling strategy, selects participating clients for the current round of training from the clients and determines the TSN transmission mode of the participating clients.

[0077] Specifically, the hierarchical scheduler converts the total federated learning training time into a joint optimization problem of user sampling + TSN scheduling, and models the joint optimization problem as a semi-Markov decision process. Then, the optimal strategies of the upper and lower layers are generated by sampling hierarchical reinforcement learning to shorten the total training data while achieving the specified accuracy as the optimization objective. An optimization problem is established, and before the start of each round of training, each edge device updates the client state through a space-time state encoder according to the current TSN network state and historical information, wherein the client state specifically includes computing power, data quality, data flow characteristics (such as data packet size, start time, transmission path, etc.) and network congestion. According to the optimization problem, the upper-layer client sampling strategy and the lower-layer injection control strategy of the hierarchical scheduler are established based on the reinforcement learning strategy and the semi-Markov decision process framework.

[0078] In addition, the embodiment needs to be explained that the specific form of the joint optimization problem of user sampling + TSN scheduling is as follows:

[0079]

[0080]

[0081] In the formula, ζ and δ represent related constants respectively; p and Δ represent the upper-layer client sampling strategy and the lower-layer injection control strategy respectively; q and G represent the proportion of client data volume and the local gradient range respectively; represents the training time of each round.

[0082] As a further scheme of the present application, the constraints of the joint optimization problem of user sampling + TSN scheduling include: for the queue e on each TSN switch, the queue length Q e,t (p, Δ) does not exceed the upper limit Q max , and the sum of all user sampling probabilities is one, and the time offset of each FL flow does not exceed the upper limit Δmax .

[0083] As a further scheme of the application, a layered reinforcement learning method is used in the joint optimization problem of user sampling + TSN scheduling, and the entire Markov decision is decomposed as follows:

[0084]

[0085] The decomposition converts the entire learning framework into a semi-Markov decision process with two different decision time ranges. Since the ultimate goal is to minimize the overall time, the upper Q function is expressed as follows:

[0086]

[0087] In the formula, γ h ∈(0, 1) represents a discount factor; and represent the start state and the start action, respectively; C r represents an instantaneous cost expression, which is specifically expressed as follows:

[0088]

[0089] The lower Q function has a similar expression form to the upper Q function, but the instantaneous cost calculation formula is expressed as follows:

[0090]

[0091] In the formula, C p represents a queue overflow penalty factor; represents the transmission time of each FL flow after a given user offset.

[0092] The central server broadcasts the current global model parameters to each participating client, and each participating client performs local training based on the received parameter information.

[0093] Specifically, at the beginning of each training round, the layered scheduler collects all client states, then generates client sampling probabilities based on the upper client sampling strategy, obtains the participating clients of this round, the central server distributes the current global model parameters to each participating client and replaces the original parameters of the local model of each participating client, each participating client collects local data using a sensor and iteratively trains the local model, then encapsulates the local model parameters after training as a time-sensitive FL flow and passes the result to the lower injection control strategy, the lower injection control strategy receives information of each participating client and collects local observation data of each participating client, then updates the state of the lower injection control strategy according to each local observation data and generates transmission decisions corresponding to each participating client.

[0094] The local parameters obtained in each round of training are encapsulated as FL flows and uploaded, and based on the FL flows, the global model parameters in the central server are updated.

[0095] Specifically, each participating client uploads the encapsulated FL flow, the hierarchical scheduler executes the selected transmission decision, injects the FL flow into the TSN network, and observes the results, including whether the injection is successful and the impact on the network state, and records the transmission time of the flow. According to the injection result, the lower layer injection control strategy obtains the corresponding reward signal, which represents the goodness of the action. The current state of each participating client, the executed action, the obtained reward, and the subsequent state are stored as experience tuples and placed in the lower layer replay buffer. After all participating clients complete the transmission in this round, the parameter aggregation is performed in the central server. Then, the total time consumption of the local training and the user with the slowest transmission time in this round is taken as the training time of this round. The upper reward signal is obtained according to the transmission time of this round and the sampling probability obtained by the upper client sampling strategy. The current state of each participating client, the executed action, the obtained reward, and the subsequent state are stored as experience tuples and placed in the upper layer replay buffer.

[0096] The training time of each round is recorded, and the results of the hierarchical scheduler are collected. The experience replay mechanism is used to optimize the decision of the hierarchical scheduler.

[0097] Specifically, it is judged whether the accumulated experience in the upper layer replay buffer and the lower layer replay buffer reaches the preset threshold. If not, the interaction collection continues. If the accumulated experience in the upper layer replay buffer and the lower layer replay buffer reaches the preset threshold, training samples are extracted from the upper layer replay buffer and the lower layer replay buffer according to the priority sampling strategy. According to the policy gradient theorem, the upper client sampling strategy parameters are updated. Then, the lower layer injection control strategy of each participating client is trained through the non-cumulative Bellman update, and the corresponding action function is established to minimize the maximum delay in a single round. Based on the action function, the lower layer injection control strategy parameters are updated. When the iteration period reaches the preset update period, the target network parameters are updated to the current estimated network parameters.

[0098] It should be further explained that the specific calculation formula for updating the upper client sampling strategy parameters is as follows:

[0099]

[0100] In the formula, α h represents the upper learning rate; Q h (s r , p r) an action value function representing the sampling strategy of the upper client, which is estimated by a deep double Q network and realizes stable learning by combining target network parameters, and which is responsible for evaluating the action p in the state s;

[0101] The specific calculation formula for updating the lower-layer injection control strategy parameters is as follows:

[0102]

[0103] In the formula, The action value function represents the action value function of the lower-layer injection control strategy.

[0104] The client selection, local training, global model parameter updating, and strategy updating are repeatedly performed until the global model reaches the specified precision or reaches the preset maximum number of iterations.

Claims

1. An adaptive hierarchical scheduling method integrating TSN and federated learning systems, characterized in that, The specific steps of this scheduling method are as follows: Ⅰ: Construct a time-sensitive industrial IoT framework and initialize the various parameters of the time-sensitive industrial IoT framework; II: The hierarchical scheduler generates the latest scheduling policy. At the beginning of each training round, based on the scheduling policy, it selects the participating clients for this training round from the clients and determines their TSN transmission mode. III: The central server broadcasts the current global model parameters to each participating client, and each participating client performs local training based on the received parameter information; IV: Encapsulate the local parameters obtained in each round of training into FL streams and upload them. Based on each FL stream, update the global model parameters in the central server. V: Record the training time for each round, collect the results of the hierarchical scheduler, and optimize the hierarchical scheduler's decisions using the experience replay mechanism; VI: Repeat the process of client selection, local training, global model parameter update, and policy update until the global model reaches the specified accuracy or the preset maximum number of iterations.

2. The adaptive hierarchical scheduling method for integrating TSN and federated learning systems according to claim 1, characterized in that, The time-sensitive industrial IoT framework described in step I specifically includes a device layer, a hierarchical scheduler, a TSN network, a TSN CQF circular queue, and a field layer; The device layer includes multiple edge devices in the Industrial Internet of Things, i.e. clients, for data collection and local computing. After the selected client completes local computing in each round, it encapsulates the transmitted data into an FL stream and transmits the FL stream to the TSN network at different times according to the decision result of the hierarchical scheduler. The hierarchical scheduler consists of a spatiotemporal state encoder, an upper-layer client sampling strategy, and a lower-layer injection control strategy, which are respectively responsible for global state modeling and joint optimization of client selection and injection time slot allocation; The spatiotemporal state encoder is used to extract information including the computing power of the federated learning client, data quality, and TSN path congestion, and generate a global state representation; the upper-layer client sampling strategy is used for the federated learning user selection strategy, which selects clients with stronger computing power and data quality that meet preset conditions to participate in this round of training; the lower-layer injection control strategy is used to optimize the real-time stream injection strategy of edge devices, that is, to inject the time slot offset of the TSN network for real-time transmission of the TSN network. The TSN network includes TSN gateways and TSN switches, which are used to receive and forward FL streams injected by edge devices and perform time-deterministic forwarding of them. The TSN gateway is used to receive FL streams and transmit them to the central server; the TSN switch is used to perform multi-hop forwarding of FL streams injected into the TSN network; and the TSN gateway and TSN switch use TSN CQF circular queues to achieve deterministic transmission of data packets. The TSN CQF circular queue consists of a send queue and a receive queue; The field layer consists of edge servers, controllers, and traditional field-level devices. The edge servers act as the central server for federated learning to aggregate parameters, while the controllers make hierarchical decisions on user selection and transmission latency.

3. The adaptive hierarchical scheduling method for integrating TSN and federated learning systems according to claim 2, characterized in that, The specific steps for the hierarchical scheduler to generate the latest scheduling policy in step II are as follows: S1.1: The hierarchical scheduler transforms the total training time of federated learning into a joint optimization problem of user sampling + TSN scheduling, and models this joint optimization problem as a semi-Markov decision process. Then, it generates the optimal policies of the upper and lower layers through sampling hierarchical reinforcement learning, with the goal of achieving the specified accuracy while reducing the total training data, and establishes the optimization problem. S1.2: Before each training round, each edge device updates its client state through a spatiotemporal state encoder based on the current TSN network state and historical information. Based on the optimization problem, and using a reinforcement learning strategy and a semi-Markov decision process framework, it establishes an upper-layer client sampling strategy and a lower-layer injection control strategy for the hierarchical scheduler.

4. The adaptive hierarchical scheduling method for integrating TSN and federated learning systems according to claim 3, characterized in that, The specific form of the joint optimization problem of user sampling + TSN scheduling described in S1.1 is as follows: In the formula, ζ and δ represent the correlation constants, respectively; p and Δ represent the upper-layer client sampling strategy and the lower-layer injection control strategy, respectively; q and G represent the proportion of client data and the local gradient range, respectively. This indicates the time allotted for each training session.

5. The adaptive hierarchical scheduling method for integrating TSN and federated learning systems according to claim 4, characterized in that, The specific steps for each participating client to perform local training based on the received parameter information, as described in step III, are as follows: S2.1: At the beginning of each training round, the hierarchical scheduler collects all client states, then generates sampling probabilities for each client based on the upper-layer client sampling strategy, obtains the participating clients for that round, and the central server sends the current global model parameters to each participating client and replaces the original parameters of the local model of each participating client. S2.2: Each participating client uses the local data collected by the sensor to iteratively train the local model, then encapsulates the local model parameters after training into a time-sensitive FL stream, and passes the result to the lower layer injection control strategy. S2.3: The lower-layer injection control strategy receives information from each participating client and collects local observation data from each participating client. Then, based on the local observation data, it updates the state of the lower-layer injection control strategy and generates the corresponding transmission decision for each participating client.

6. The adaptive hierarchical scheduling method for integrating TSN and federated learning systems according to claim 5, characterized in that, The specific steps for updating the global model parameters in the central server are as follows: S3.1: Each participating client uploads the encapsulated FL stream. The hierarchical scheduler executes the selected transmission decision, injects the FL stream into the TSN network, and observes the results. The observation results include whether the injection was successful and the impact on the network status. At the same time, the transmission time of the stream is recorded. S3.2: Based on the injection result, the lower-level injection control strategy obtains the corresponding reward signal. This reward signal represents the quality of the action. The current state of each participating client, the action executed, the reward obtained, and the subsequent state are stored as an experience tuple and placed into the lower-level replay buffer. S3.3: After all participating clients have completed the transmission, the parameters are aggregated on the central server. Then, the total time spent on local training and the transmission of the slowest user in this round is used as the training time for this round. The upper-layer reward signal is obtained based on the transmission time of this round and the sampling probability obtained by the upper-layer client sampling strategy. The current state, the action performed, the reward obtained and the subsequent state of each participating client are stored as experience tuples and put into the upper-layer replay buffer.

7. The adaptive hierarchical scheduling method for integrating TSN and federated learning systems according to claim 6, characterized in that, The specific steps for optimizing the hierarchical scheduler decision-making described in step V are as follows: S4.1: Determine whether the upper and lower playback buffers have accumulated enough experience to reach the preset threshold. If not, continue interactive acquisition. S4.2: If the accumulated experience in the upper replay buffer and the lower replay buffer reaches the preset threshold, then training samples are extracted from the upper replay buffer and the lower replay buffer respectively according to the priority sampling strategy. S4.3: According to the policy gradient theorem, update the upper-layer client sampling policy parameters, and then train the lower-layer injection control policy of each participating client through non-cumulative Bellman update. With the goal of minimizing the maximum latency in a single round, establish the corresponding action function, and update the lower-layer injection control policy parameters based on the action function. S4.4: When the iteration period reaches the preset update period, the target network parameters are updated synchronously to the current estimated network parameters.

8. The adaptive hierarchical scheduling method for integrating TSN and federated learning systems according to claim 7, characterized in that, The specific calculation formula for updating the upper-layer client sampling strategy parameters as described in S4.3 is as follows: In the formula, α h Q represents the upper-level learning rate; h (s r ,p r ) represents the action value function of the upper-layer client sampling strategy. It is estimated using a deep dual Q network and combined with the target network parameters to achieve stable learning. This action value function is responsible for evaluating the action p in state s. The specific calculation formula for updating the lower-level injection control strategy parameters as described in S4.3 is as follows: In the formula, The action value function represents the lower-level injection control strategy.

Citation Information

Patent Citations

  • 5G and TSN joint scheduling method based on wireless channel information

    CN113630893A