Clustering network flow scheduling method, device and equipment based on star flash
By constructing a traffic scheduling model for the StarSpark cluster network and a Markov decision process involving multiple management nodes, the traffic scheduling problem of the cluster network in the StarSpark system was solved, achieving high reliability, low latency communication quality, and load balancing.
Patent Information
- Application Number
- CN202511198053.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-14
AI Technical Summary
Existing short-range wireless communication technologies such as Bluetooth and WiFi are insufficient to meet the requirements of wide coverage, high reliability, and low latency in industrial scenarios. Furthermore, the traffic scheduling strategy of the Starflash system in clustered networks has not been fully studied, resulting in insufficient service flow scheduling performance.
We design a clustered network traffic scheduling method based on star-flash. By constructing a traffic scheduling model, we introduce a Markov decision process with multi-management node cooperation and a reinforcement learning algorithm to optimize resource allocation and achieve load balancing and low-latency transmission.
It improves the scheduling performance of service flows in clustered networks, meets the goals of network load balancing and minimizing end-to-end communication latency, and achieves high reliability and low latency communication quality.
Smart Images

Figure CN120957149A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial network technology, and in particular to a clustered network traffic scheduling method, apparatus, and equipment based on star-flash. Background Technology
[0002] Short-range wireless communication is a key technology for building smart environments. Traditional short-range wireless communication technologies, such as Bluetooth and WiFi, suffer from limited coverage, low accuracy, long latency, and poor anti-interference capabilities, making it difficult to meet the stringent requirements of emerging industrial scenarios for wide coverage, high reliability, and low latency. StarSpeed technology, as an emerging short-range wireless communication technology, optimizes the transmission distance, latency, and anti-interference performance of wireless communication systems by introducing new features such as 5G Polar coding, ultra-short frames, and frequency hopping signal splicing, solving the industry pain points faced by Bluetooth and WiFi. However, StarSpeed only defines the basic architecture and node communication model. In practical clustered industrial network periodic traffic scheduling scenarios, the scheduling strategy for service flows in the StarSpeed system requires further research. Combining the multi-agent deep reinforcement learning technology emerging in distributed network scenarios, it is necessary to propose a StarSpeed scheduling method to achieve more intelligent and dynamic resource allocation and improve the scheduling performance of service flows in clustered StarSpeed networks. Summary of the Invention
[0003] The purpose of this invention is to provide a method, apparatus, and device for traffic scheduling in clustered networks based on StarSpark, specifically designed for traffic scheduling in clustered network mode within a StarSpark system. First, a StarSpark-based clustered network traffic scheduling model is designed to support highly reliable and low-latency transmission of cluster node data within the StarSpark system. Then, based on the model design, a large number of constraint optimization problems related to control flow during scheduling are introduced. To improve the efficiency of solving these constraint optimization problems, a multi-management node collaborative traffic scheduling algorithm is further designed. This algorithm can rationally plan subframes and time slots in the StarSpark clustered network system, optimizing traffic scheduling for nodes within each cluster and across cluster management nodes. This enables the system to improve load balancing capabilities while reducing end-to-end transmission latency of service flows.
[0004] This invention provides a clustered network traffic scheduling method based on star flash, comprising: Collect network status information and service flow information of nodes within the cluster; Construct a clustered network traffic time slot scheduling model based on star flash; Based on the optimization problem of the star-flash clustered network traffic scheduling model, a network traffic scheduling constrained optimization problem is established with the objectives of network load balancing and minimizing end-to-end latency. Based on the constrained optimization problem, a Markov decision process with multi-management node collaboration is constructed. The Markov decision process is solved using a traffic scheduling algorithm based on multi-management node collaboration. The optimal resource allocation result is obtained by running a multi-management node collaborative traffic scheduling algorithm, and traffic scheduling is performed on the star-flash cluster network based on the resource allocation result.
[0005] Preferably, the collection of network status information and service flow information of nodes within the cluster includes: Each cluster head node within a sub-cluster domain is defined as the management G node for that domain. The management G node is responsible for managing and scheduling the data transmission of terminal T nodes within its domain and interacting with other cluster head G nodes. The management G node collects link status information of terminal T nodes within its cluster in real time, including at least link connectivity and bandwidth utilization. The terminal T nodes actively report service flow characteristic information to the management G node of their respective cluster. The service flow characteristic information includes at least the service flow ID, source address, destination address, data packet size, service flow cycle, and latency requirements.
[0006] Preferably, the construction of the clustered network traffic time slot scheduling model based on star flash includes: A star-studded superframe structure is constructed, defining the duration of the star-studded superframe as the least common multiple of the periods of all periodic control flows, and constraining the length of the superframe to be N times the length of the subframe, where N is any value from a preset set of integers. The superframe length h constructed by the time slot scheduling model is... p The formula is described as follows: , Among them, the function Used to calculate the least common multiple. The length represents the period of each control flow, F is the set of control flows, and length is... frame It is the length of the subframe; The subframes within the superframe are functionally divided into cluster scheduling subframes and cross-cluster scheduling subframes. Cluster scheduling subframes are used to orchestrate traffic transmission between T nodes and G nodes within a cluster, while cross-cluster scheduling subframes are used to orchestrate traffic interaction between G nodes in different clusters. The clustered scheduling subframes and cross-cluster scheduling subframes are divided into protection intervals between uplink transmission intervals, downlink transmission intervals, and uplink / downlink time slot intervals. Construct a traffic transmission model during the service flow scheduling process. When traffic is scheduled to the m-th star node, the data rate of that node transmitting to the destination star node d is... Described as: in, It is an indicator symbol, if the m-th node's... If the link is an output link, then... If the m-th node's... If the link is an input link, then... ,if If it is a link that is not a node m, then ; In the link The data rate of all service flows with destination address d node.
[0007] Preferably, the optimization problem based on the star-flash clustering network traffic scheduling model includes: Construct a multi-objective optimization function that includes a load balancing sub-objective and a low latency sub-objective, and the optimization function is described as follows: , in, It is the bandwidth resource utilization rate of each link in the uplink and downlink transmission intervals, delay f It is the end-to-end latency of each business flow. It is the weighting factor of the sub-objective; Based on the principle of the star-flash clustered network traffic scheduling model, constraints are established for the optimization problem, and these constraints are configured as follows: The scheduling cycle constraint is configured to allow traffic transmission within only one scheduling cycle. The interval capacity constraint is configured such that the uplink and downlink data transmitted in each subframe interval must not exceed the maximum data rate capacity of that interval. The formula is described as follows: in, This indicates the size of the service flow data packets, and "band" indicates the link bandwidth. It refers to the transmit power on the link. It is the link gain. Interference caused by data transmission through other links, length slot It is the length of the transmission interval, and slot represents the specific transmission interval; Traffic transmission constraints are configured such that the data rate leaving the network from the destination strobe node should be equal to the sum of the data rates sent from all nodes in the network to the destination strobe node, i.e.: , in, It is the data rate leaving the destination star node, and M represents the set of all network nodes that send data streams to node d.
[0008] Preferably, the construction of a Markov decision process for multi-management node collaboration based on the constrained optimization problem includes: Define the state space, action space, and reward function of a Markov decision process; wherein each cluster head node G observes the traffic characteristics and link states of its respective cluster terminal nodes T, and the traffic characteristics include at least the traffic size T. u (t), period hp u (t), address information, priority pri u (t), Link bandwidth resource utilization As the state space s t The elements, described by the formula: Action space a t The elements include subframe clustering allocation, transmission time slot interval allocation, bandwidth allocation, and forwarding path allocation, described by the formula: , Among them, subframe clustering allocation This represents a communication allocation subframe for nodes within a cluster, meaning that only nodes within the specified cluster are allowed to communicate with each other within the duration of this subframe. `slot(t)` represents the allocated transmission time slot interval, dividing it into uplink / downlink / guard time slot intervals; `band(t)` represents the bandwidth allocated to users transmitting data within the corresponding interval; `route(t)` represents the forwarding path planned for the service flow; and `src`... u (t) represents the source address, dst u (t) represents the destination address; Cooperative reward function r co The load balancing degree and end-to-end latency of the service flow scheduling are positively correlated, as described by the formula: in, Discount factor .
[0009] Preferably, the traffic scheduling algorithm based on multi-management node cooperation for solving the Markov decision process includes: Each cluster head node G is defined as an independent intelligent agent, and the intelligent agent maintains a strategy estimation neural network and a state evaluation neural network. After each iteration and state transition, the neural network is trained. The cluster head G node collects four-tuples from the experience pool in batches to form trajectory sequences. At the same time, it interacts with every other cluster head G node in the environment (excluding itself) to obtain trajectory sequences from the experience pool. Gradient ascent is used to maximize the target. Used to update the parameters of the policy estimation neural network. The formula is described as follows: Where, parameter T is the length of the trajectory sequence, and t0 is the starting step of the trajectory sequence. It is a clipping function used to restrict the range of values to a certain value. Within the range, the original parameters used in the previous training step of the policy estimation neural network are represented as follows: The new parameters used in the current step are represented as a(t) represents the action chosen at time step t, and s(t) represents the state at time step t. Indicates in the parameter The probability of choosing action a(t) given a state s(t) by a neural network is estimated by the following policy. Indicates the original parameters The probability of choosing action a(t) given a state s(t) by a neural network is estimated by the following policy. Let A(t) be a hyperparameter, and its value be described as the dominant function value: , in, Here, t0 is the discount factor, t0 is the starting step of the trajectory, T is the total length of the trajectory, and t0+T−1−t is the number of steps from time t to the end of the trajectory. Indicates time step Timing difference error, timing difference error r(t) represents the instantaneous reward obtained at time step t, and w is the discount factor. This represents the estimated state value function for the next state. This represents the estimated value of the state value function for the current state. Update the parameters of the state evaluation neural network The loss needs to be minimized using gradient descent; the loss function is... Described as: , It is the state value function of state s(t) under policy μ0. It is the state value function of state s(t) under policy μ.
[0010] Preferably, the policy estimation neural network is used to calculate the action corresponding to the current state based on the current policy, and the state evaluation neural network is used to evaluate the value of the current state; both neural networks have an input layer, an intermediate layer and an output layer; the dimension of the input layer is equal to the dimension of the partially observable state space, the intermediate layer contains a K-layer fully connected neural network, the dimension of the output layer of the policy estimation neural network is equal to the dimension of the action space, and the dimension of the output layer of the state evaluation neural network is equal to 1.
[0011] The present invention also provides a clustered network traffic scheduling device based on star flash, comprising: The cluster node status and service flow information collection module is used to collect network status information and service flow information of nodes within the cluster; The module for constructing a clustered network traffic scheduling model based on StarSpark is used to build the frame structure and traffic transmission model for traffic scheduling in the StarSpark clustered network system. The optimization problem construction module based on the star-flash clustered network traffic scheduling model is used to establish a network traffic scheduling constrained optimization problem with the objectives of network load balancing and minimizing end-to-end latency. A Markov decision process construction module for multi-management node collaboration is used to construct a Markov decision process for multi-management node collaboration based on the constrained optimization problem. A multi-management-node collaborative traffic scheduling algorithm design module is used to solve the periodic control traffic within and between clusters in the Markov decision process based on the multi-management-node collaborative traffic scheduling algorithm. The solution module is used to run a traffic scheduling algorithm with multiple management nodes to solve for the optimal resource allocation result, and to perform traffic scheduling on the star-flash cluster network based on the resource allocation result.
[0012] The present invention also provides an electronic device, comprising: The memory is used to store the processing program; The processor, when executing the processing program, implements the clustered network traffic scheduling method based on star flash as described in the embodiments of the present invention.
[0013] The present invention also provides a readable storage medium storing a processing program, which, when executed by a processor, implements the clustered network traffic scheduling method based on star flash as described in the embodiments of the present invention.
[0014] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a clustered network traffic scheduling method based on star-flash. By constructing a clustered network traffic scheduling model based on star-flash, it can ensure high reliability and low latency communication quality during the clustered network traffic scheduling process. On this basis, a traffic scheduling algorithm with multi-management node cooperation is designed, which can further improve the scheduling performance of service flows in the distributed network and meet the goals of network load balancing and minimizing end-to-end communication latency. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the clustered network traffic scheduling method based on star flash in an embodiment of the present invention. Figure 2 This is a flowchart of a clustered network traffic scheduling method based on star flash according to another embodiment of the present invention; Figure 3This is a structural block diagram of a clustered network traffic scheduling device based on star flash according to an embodiment of the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] The term "comprising" and its variations as used herein are open-ended inclusion, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0018] Figure 1 This paper describes an application scenario of an embodiment of the present invention. The network topology in this scenario is divided into four clusters. Each cluster includes a management node G (cluster head) and cluster member nodes T. Nodes communicate with each other over short distances of 10-20 meters. The transmit power is set to 10 dBm, the channel frequency band is 5 GHz to support high-speed data transmission, and the frequency bandwidth is 20 MHz. A barrier separates cluster 1 and cluster 2, preventing direct communication between the two clusters. Communication between the two clusters requires multi-hop relay nodes. Each T node periodically sends control-type service flows, with data packet sizes ranging from 50 bytes to 1 KB. To optimize network traffic scheduling performance, the time slot intervals, paths, and bandwidth required for data forwarding are determined by a traffic scheduling algorithm involving multiple management nodes, improving the system's load balancing capability and optimizing end-to-end latency of service flow transmission. The algorithm uses the Adam optimizer with a learning rate of 1e-4. The intermediate layers consist of two fully connected neural networks. The experience pool size is set to 1024, the sampling batch size is set to 32, and the discount factor is 0.95 for all layers.
[0019] To improve the scheduling efficiency of a star-flash clustering network system, this application provides a star-flash-based clustering network traffic scheduling method. For example... Figure 2 As shown, the scheduling method includes the following steps: S101: Collect network status information and service flow information of nodes within the cluster; Step S101: Cluster node status and service flow information collection; Specifically, the cluster head node in each sub-cluster domain is defined as the management G node of that domain, responsible for managing and scheduling the data transmission of the terminal T nodes in that domain, and completing the interaction with other cluster head G nodes; The management G node of the cluster head collects the link connectivity, bandwidth and other statuses of the terminal T nodes in the cluster in real time; The terminal T nodes upload the characteristic information of the service flow to the management G node of their respective cluster, including the service flow ID, source address, destination address, data packet size, service flow cycle, latency requirements, etc.
[0020] S102: Construct a clustered network traffic slot scheduling model based on star-flash; specifically, the model sets the least common multiple of the periodic control traffic period as the duration of the star-flash superframe, and requires the superframe length to be equal to N times the subframe length, where N=6,12,18,…,48. The superframe length h constructed by the slot scheduling model is... p The formula is described as follows: , Among them, the function Used to calculate the least common multiple. The length represents the period of each control flow, F is the set of control flows, and length is... frame It is the length of the subframe; Each superframe is considered a complete traffic scheduling cycle. The function of each subframe within the superframe is defined, including clustered scheduling of subframes. Cross-cluster scheduling subframes Cluster scheduling subframes are used to orchestrate traffic transmission between nodes T and G within a cluster, denoted as . Cross-cluster scheduling subframes are used to orchestrate traffic interactions between G nodes in different clusters, denoted as... ,and .in, This represents a cluster.
[0021] The subframes within the superframe are functionally divided into cluster-based scheduling subframes and cross-cluster scheduling subframes. Cluster-based scheduling subframes are used to orchestrate traffic transmission between T nodes and G nodes within a cluster, while cross-cluster scheduling subframes are used to orchestrate traffic interaction between G nodes in different clusters. The transmission range of a cluster-based scheduling subframe includes the uplink transmission range from the T node to the G node. Downlink transmission interval from G node to T node and the protection interval between uplink and downlink time slots. The uplink transmission interval from node T to node G is used to send data from node T to node G, and the downlink transmission interval from node G to node T is used to send data from node G to node T. The protection interval between the uplink and downlink time slots is used to avoid conflicts between the uplink and downlink time slots. For cross-cluster scheduling subframes, since one of the G nodes in a cluster switches to a T node when communicating between G nodes in two clusters, the transmission interval planning for cross-cluster scheduling subframes also includes uplink transmission intervals, downlink transmission intervals, and protection intervals.
[0022] The clustered scheduling subframes and cross-cluster scheduling subframes are divided into protection intervals between uplink transmission intervals, downlink transmission intervals, and uplink / downlink time slot intervals. Construct a traffic transmission model during the service flow scheduling process. When traffic is scheduled to the m-th star node, the data rate of that node transmitting to the destination star node d is... Described as: in, It is an indicator symbol, if the m-th node's... If the link is an output link, then... If the m-th node's... If the link is an input link, then... ,if If it is a link that is not a node m, then ; In the link The data rate of all service flows with destination address d node.
[0023] S103: Based on the star-flash clustered network traffic scheduling model, construct an optimization problem for network traffic scheduling with the objectives of network load balancing and minimizing end-to-end latency; Specifically, the optimization objectives of this problem include two sub-objectives: improving the load balancing capability of the star-flash clustering network system and reducing the end-to-end transmission latency of service flows. Based on the designed sub-objectives, the optimization function is described as follows: , in, It is the bandwidth resource utilization rate of each link in the uplink and downlink transmission intervals, delay f It is the end-to-end latency of each business flow. It is the weighting factor of the sub-objective; Based on the principle of the star-flash clustered network traffic scheduling model, constraints are established for the optimization problem, and these constraints are configured as follows: Scheduling cycle constraints: Due to the periodic nature of the control flow, the system is configured to transmit traffic within only one scheduling cycle. The interval capacity constraint requires that the uplink and downlink data transmitted in each subframe interval must not exceed the maximum data rate capacity of that interval. The configuration to ensure that the uplink and downlink data transmitted in each subframe interval does not exceed the maximum data rate capacity of that interval is described by the formula: in, This indicates the size of the service flow data packets, and "band" indicates the link bandwidth. It refers to the transmit power on the link. It is the link gain. Interference caused by data transmission through other links, length slot It is the length of the transmission interval, and slot represents the specific transmission interval; Traffic transmission constraints are configured such that the data rate leaving the network from the destination strobe node should be equal to the sum of the data rates sent from all nodes in the network to the destination strobe node, i.e.: , in, It is the data rate leaving the destination star node, and M represents the set of all network nodes that send data streams to node d.
[0024] S104: Construct a Markov decision process for multi-management node collaboration based on the constrained optimization problem; specifically, based on the objective and constraints of the optimization problem, design a partially observable state space, action space, and cooperative reward for the Markov decision process.
[0025] Define the state space, action space, and reward function of a Markov decision process; wherein each cluster head node G observes the traffic characteristics and link states of its respective cluster terminal nodes T, and the traffic characteristics include at least the traffic size T. u (t), period hp u (t), address information, priority pri u (t), Link bandwidth resource utilization As the state space s t The elements, described by the formula: Action space a t The elements include subframe clustering allocation, transmission time slot interval allocation, bandwidth allocation, and forwarding path allocation, described by the formula: , Among them, subframe clustering allocation This represents a communication allocation subframe for nodes within a cluster, meaning that only nodes within the specified cluster are allowed to communicate with each other within the duration of this subframe. `slot(t)` represents the allocated transmission time slot interval, dividing it into uplink / downlink / guard time slot intervals; `band(t)` represents the bandwidth allocated to users transmitting data within the corresponding interval; `route(t)` represents the forwarding path planned for the service flow; and `src`... u (t) represents the source address, dst u (t) represents the destination address; Cooperative reward function r co The load balancing degree and end-to-end latency of the service flow scheduling are positively correlated, as described by the formula: in, Discount factor .
[0026] S105: Solve the Markov decision process using a traffic scheduling algorithm based on multi-management node cooperation; Specifically, to efficiently solve the optimization problem based on the star-flash clustering network traffic scheduling model, a traffic scheduling algorithm based on multi-agent near-end policy optimization is designed. This algorithm defines each cluster head node G as an agent, which maintains two neural networks: a policy estimation neural network and a state evaluation neural network. The policy estimation neural network calculates the action corresponding to the current state based on the current policy, while the state evaluation neural network evaluates the value of the current state. Both neural networks consist of an input layer, intermediate layers, and an output layer. The dimension of the input layer is equal to the dimension of the partially observable state space, the intermediate layers contain K fully connected neural networks, the output layer dimension of the policy estimation neural network is equal to the dimension of the action space, and the output layer dimension of the state evaluation neural network is 1.
[0027] During the state transition process of this algorithm, each cluster head node G observes the collected states at each iteration step. The policy estimation neural network calculates the action based on the state-action mapping policy and issues the corresponding decision. After the environment executes the action instruction, the G node receives the corresponding reward value and transitions to the next state. If the resource utilization of the next state exceeds the capacity of the link slot or bandwidth, or does not meet the end-to-end latency requirements of the service flow, the scheduling fails, and the algorithm starts a new round of iteration; otherwise, it continues to iterate in the current round. The G node stores the four-tuple of current state, current action, reward, and next state obtained at each step into the experience pool.
[0028] Each cluster head node G is defined as an independent intelligent agent, and the intelligent agent maintains a strategy estimation neural network and a state evaluation neural network. After each iteration and state transition, the neural network is trained. Specifically, the cluster head G node collects a batch of quadruplets from the experience pool to form a trajectory sequence, and simultaneously interacts with every other cluster head G node in the environment (excluding itself) to obtain trajectory sequences from the experience pool. Gradient ascent is used to maximize the target. Used to update the parameters of the policy estimation neural network. The formula is described as follows: Where, parameter T is the length of the trajectory sequence, and t0 is the starting step of the trajectory sequence. It is a clipping function used to restrict the range of values to a certain value. Within the range, the original parameters used in the previous training step of the policy estimation neural network are represented as follows: The new parameters used in the current step are represented as a(t) represents the action chosen at time step t, and s(t) represents the state at time step t. Indicates in the parameter The probability of choosing action a(t) given a state s(t) by a neural network is estimated by the following policy. Indicates the original parameters The probability of choosing action a(t) given a state s(t) by a neural network is estimated by the following policy. Let A(t) be a hyperparameter, and its value be described as the dominant function value: , in, Here, t0 is the discount factor, t0 is the starting step of the trajectory, T is the total length of the trajectory, and t0+T−1−t is the number of steps from time t to the end of the trajectory. Indicates time step Timing difference error, timing difference error r(t) represents the instantaneous reward obtained at time step t, and w is the discount factor. This represents the estimated state value function for the next state. This represents the estimated value of the state value function for the current state. Update the parameters of the state evaluation neural network The loss needs to be minimized using gradient descent; the loss function is... Described as: , It is the state value function of state s(t) under policy μ0. It is the state value function of state s(t) under policy μ.
[0029] S106: Run a multi-management node collaborative traffic scheduling algorithm to solve the problem, obtain the optimal resource allocation result, and perform traffic scheduling on the Starflash cluster network according to the resource allocation result.
[0030] The above scheme manages G nodes to collect link status (on / off / bandwidth) in real time and service flow characteristics (latency requirements, data volume) reported by terminal T nodes, enabling scheduling decisions to respond quickly to dynamic network changes. It is particularly suitable for industrial scenarios with deterministic latency requirements. The superframe structure sets the least common multiple of the periodic control flow period as the superframe duration and divides it into clustered scheduling subframes (intra-cluster TG communication) and cross-cluster scheduling subframes (GG interaction). Combined with uplink and downlink transmission intervals and protection intervals, it effectively avoids conflicts and reduces packet loss rate. With link bandwidth utilization (U(l)) and end-to-end service flow latency (D(f)) as joint optimization objectives, it flexibly adapts to different scenario requirements (such as high real-time priority or high throughput priority) through weight factors α / β. Through scheduling cycle constraints (single cycle planning), interval capacity constraints (subframe data volume ≤ maximum rate) and flow conservation constraints (destination node output = total network output) To ensure the feasibility and fairness of resource allocation, a Markov Decision Process (MDP) model is used to model complex scheduling problems, abstracting them into a state space (traffic characteristics + bandwidth utilization), an action space (subframe allocation / time slot planning / bandwidth allocation / path selection), and a reward function (load balancing + latency optimization), adapting to reinforcement learning frameworks. A multi-agent collaboration mechanism is adopted: each cluster head G node acts as an independent agent, generating actions through a policy estimation neural network (actor) and evaluating values through a state evaluation neural network (commentator). Combined with experience pool replay and cross-agent experience sharing, convergence to the global optimum is accelerated. A cluster management mechanism is used: the management G node is only responsible for scheduling terminal T nodes within its own cluster. Cross-cluster interaction is completed through dedicated cross-cluster scheduling subframes, avoiding the single-point bottleneck of centralized control. When adding a new cluster, only the local G node parameters need to be configured, without global reconstruction, making it suitable for large-scale industrial IoT scenarios. By combining the characteristics of the starburst physical layer, cluster architecture, and multi-agent reinforcement learning, the bottlenecks of coverage, latency, and reliability of traditional short-range communication in industrial scenarios are solved, achieving deterministic low latency (microsecond level), high reliability transmission (anti-interference + protection mechanism), dynamic load balancing (multi-objective optimization + reinforcement learning), and flexible scalability (distributed architecture).
[0031] Example 2 Based on the same concept, such as Figure 3 As shown, the present invention provides a clustered network traffic scheduling device based on star flash, comprising: The cluster node status and service flow information collection module is used to collect network status information and service flow information of nodes within the cluster; The module for constructing a clustered network traffic scheduling model based on StarSpark is used to build the frame structure and traffic transmission model for traffic scheduling in the StarSpark clustered network system. The optimization problem construction module based on the star-flash clustered network traffic scheduling model is used to establish a network traffic scheduling constrained optimization problem with the objectives of network load balancing and minimizing end-to-end latency. A Markov decision process construction module for multi-management node collaboration is used to construct a Markov decision process for multi-management node collaboration based on the constrained optimization problem. A multi-management-node collaborative traffic scheduling algorithm design module is used to solve the periodic control traffic within and between clusters in the Markov decision process based on the multi-management-node collaborative traffic scheduling algorithm. The solution module is used to run a traffic scheduling algorithm with multiple management nodes to solve the problem, obtain the optimal resource allocation result, and perform traffic scheduling on the star-flash cluster network according to the resource allocation result. The implementation principle of the above module has been described in the previous embodiments, so it will not be repeated here.
[0032] Example 3 Based on the same concept, an electronic device is also provided in some embodiments of this application. This electronic device includes a memory and a processor, wherein the memory stores a processing program, and the processor executes the processing program according to instructions. When the processor executes the processing program, the star-flash-based clustered network traffic scheduling method described in the foregoing embodiments is implemented.
[0033] In some embodiments of this application, a readable storage medium is also provided. This readable storage medium can be a non-volatile readable storage medium or a volatile readable storage medium. The readable storage medium stores instructions that, when executed on a computer, cause an electronic device containing this readable storage medium to perform the aforementioned star-flash-based clustered network traffic scheduling method.
[0034] It is understood that, for the aforementioned clustered network traffic scheduling methods based on star-flash technology, if all are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0035] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0036] The program code for executing the technical solutions disclosed in this application can be written in any combination of one or more programming languages. These programming languages include object-oriented programming languages—such as Java and C++—and conventional procedural programming languages—such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A clustered network traffic scheduling method based on star flash, characterized in that, include: Collect network status information and service flow information of nodes within the cluster; Construct a clustered network traffic time slot scheduling model based on star flash; Based on the optimization problem of the star-flash clustered network traffic scheduling model, a network traffic scheduling constrained optimization problem is established with the objectives of network load balancing and minimizing end-to-end latency. Based on the constrained optimization problem, a Markov decision process with multi-management node collaboration is constructed. The Markov decision process is solved using a traffic scheduling algorithm based on multi-management node collaboration. The optimal resource allocation result is obtained by running a multi-management node collaborative traffic scheduling algorithm, and traffic scheduling is performed on the star-flash cluster network based on the resource allocation result.
2. The clustered network traffic scheduling method based on star flash as described in claim 1, characterized in that, The collection of network status information and service flow information of nodes within the cluster includes: Each cluster head node within a sub-cluster domain is defined as the management G node for that domain. The management G node is responsible for managing and scheduling the data transmission of terminal T nodes within its domain and interacting with other cluster head G nodes. The management G node collects link status information of terminal T nodes within its cluster in real time, including at least link connectivity and bandwidth utilization. The terminal T nodes actively report service flow characteristic information to the management G node of their respective cluster. The service flow characteristic information includes at least the service flow ID, source address, destination address, data packet size, service flow cycle, and latency requirements.
3. The clustered network traffic scheduling method based on star flash as described in claim 1, characterized in that, The construction of the clustered network traffic time slot scheduling model based on star flash includes: A star-studded superframe structure is constructed, defining the duration of the star-studded superframe as the least common multiple of the periods of all periodic control flows, and constraining the length of the superframe to be N times the length of the subframe, where N is any value from a preset set of integers. The superframe length h constructed by the time slot scheduling model is... p The formula is described as follows: , Among them, the function Used to calculate the least common multiple. The length represents the period of each control flow, F is the set of control flows, and length is... frame It is the length of the subframe; The subframes within the superframe are functionally divided into cluster scheduling subframes and cross-cluster scheduling subframes. Cluster scheduling subframes are used to orchestrate traffic transmission between T nodes and G nodes within a cluster, while cross-cluster scheduling subframes are used to orchestrate traffic interaction between G nodes in different clusters. The clustered scheduling subframes and cross-cluster scheduling subframes are divided into protection intervals between uplink transmission intervals, downlink transmission intervals, and uplink / downlink time slot intervals. Construct a traffic transmission model during the service flow scheduling process. When traffic is scheduled to the m-th star node, the data rate of that node transmitting to the destination star node d is... Described as: in, It is an indicator symbol, if the m-th node's... If the link is an output link, then... If the m-th node's... If the link is an input link, then... ,if If it is a link that is not a node m, then ; In the link The data rate of all service flows with destination address d node.
4. The clustered network traffic scheduling method based on star flash as described in claim 1, characterized in that, The optimization problem based on the star-flash clustering network traffic scheduling model includes: Construct a multi-objective optimization function that includes a load balancing sub-objective and a low latency sub-objective, and the optimization function is described as follows: , in, It is the bandwidth resource utilization rate of each link in the uplink and downlink transmission intervals, delay f It is the end-to-end latency of each business flow. It is the weighting factor of the sub-objective; Based on the principle of the star-flash clustered network traffic scheduling model, constraints are established for the optimization problem, and these constraints are configured as follows: The scheduling cycle constraint is configured to allow traffic transmission within only one scheduling cycle. The interval capacity constraint is configured such that the uplink and downlink data transmitted in each subframe interval must not exceed the maximum data rate capacity of that interval. The formula is described as follows: in, This indicates the size of the service flow data packets, and "band" indicates the link bandwidth. It refers to the transmit power on the link. It is the link gain. Interference caused by data transmission through other links, length slot It is the length of the transmission interval, and slot represents the specific transmission interval; Traffic transmission constraints are configured such that the data rate leaving the network from the destination strobe node should be equal to the sum of the data rates sent from all nodes in the network to the destination strobe node, i.e.: , in, It is the data rate leaving the destination star node, and M represents the set of all network nodes that send data streams to node d.
5. The clustered network traffic scheduling method based on star flash as described in claim 1, characterized in that, The Markov decision process for constructing multi-management node collaboration based on the constrained optimization problem includes: Define the state space, action space, and reward function of a Markov decision process; wherein each cluster head node G observes the traffic characteristics and link states of its respective cluster terminal nodes T, and the traffic characteristics include at least the traffic size T. u (t), period hp u (t), address information, priority pri u (t), Link bandwidth resource utilization As the state space s t The elements, described by the formula: Action space a t The elements include subframe clustering allocation, transmission time slot interval allocation, bandwidth allocation, and forwarding path allocation, described by the formula: , Among them, subframe clustering allocation This represents a communication allocation subframe for nodes within a cluster, meaning that only nodes within the specified cluster are allowed to communicate with each other within the duration of this subframe. `slot(t)` represents the allocated transmission time slot interval, dividing it into uplink / downlink / guard time slot intervals; `band(t)` represents the bandwidth allocated to users transmitting data within the corresponding interval; `route(t)` represents the forwarding path planned for the service flow; and `src`... u (t) represents the source address, dst u (t) represents the destination address; Cooperative reward function r co The load balancing degree and end-to-end latency of the service flow scheduling are positively correlated, as described by the formula: in, Discount factor .
6. The clustered network traffic scheduling method based on star flash as described in claim 1, characterized in that, The traffic scheduling algorithm based on multi-management node cooperation solves the Markov decision process by including: Each cluster head node G is defined as an independent intelligent agent, and the intelligent agent maintains a strategy estimation neural network and a state evaluation neural network. After each iteration and state transition, the neural network is trained. The cluster head G node collects four-tuples from the experience pool in batches to form trajectory sequences. At the same time, it interacts with every other cluster head G node in the environment (excluding itself) to obtain trajectory sequences from the experience pool. Gradient ascent is used to maximize the target. Used to update the parameters of the policy estimation neural network. The formula is described as follows: Where, parameter T is the length of the trajectory sequence, and t0 is the starting step of the trajectory sequence. It is a clipping function used to restrict the range of values to a certain value. Within the range, the original parameters used in the previous training step of the policy estimation neural network are represented as follows: The new parameters used in the current step are represented as a(t) represents the action chosen at time step t, and s(t) represents the state at time step t. Indicates in the parameter The probability of choosing action a(t) given a state s(t) by a neural network is estimated by the following policy. Indicates the original parameters The probability of choosing action a(t) given a state s(t) by a neural network is estimated by the following policy. Let A(t) be a hyperparameter, and its value be described as the dominant function value: , in, Here, t0 is the discount factor, t0 is the starting step of the trajectory, T is the total length of the trajectory, and t0+T−1−t is the number of steps from time t to the end of the trajectory. Indicates time step Timing difference error, timing difference error r(t) represents the instantaneous reward obtained at time step t, and w is the discount factor. This represents the estimated state value function for the next state. This represents the estimated value of the state value function for the current state. Update the parameters of the state evaluation neural network The loss needs to be minimized using gradient descent; the loss function is... Described as: , It is the state value function of state s(t) under policy μ0. It is the state value function of state s(t) under policy μ.
7. The clustered network traffic scheduling method based on star flash as described in claim 6, characterized in that, The policy estimation neural network is used to calculate the action corresponding to the current state based on the current policy, and the state evaluation neural network is used to evaluate the value of the current state. Both neural networks have an input layer, an intermediate layer, and an output layer. The input layer dimension is equal to the dimension of the partially observable state space, and the intermediate layer contains a K-layer fully connected neural network. The output layer dimension of the policy estimation neural network is equal to the dimension of the action space, and the output layer dimension of the state evaluation neural network is equal to 1.
8. A clustered network traffic scheduling device based on star flash, characterized in that, include: The cluster node status and service flow information collection module is used to collect network status information and service flow information of nodes within the cluster; The module for constructing a clustered network traffic scheduling model based on StarSpark is used to build the frame structure and traffic transmission model for traffic scheduling in the StarSpark clustered network system. The optimization problem construction module based on the star-flash clustered network traffic scheduling model is used to establish a network traffic scheduling constrained optimization problem with the objectives of network load balancing and minimizing end-to-end latency. A Markov decision process construction module for multi-management node collaboration is used to construct a Markov decision process for multi-management node collaboration based on the constrained optimization problem. A multi-management-node collaborative traffic scheduling algorithm design module is used to solve the periodic control traffic within and between clusters in the Markov decision process based on the multi-management-node collaborative traffic scheduling algorithm. The solution module is used to run a traffic scheduling algorithm with multiple management nodes to solve for the optimal resource allocation result, and to perform traffic scheduling on the star-flash cluster network based on the resource allocation result.
9. An electronic device, characterized in that, include: The memory is used to store the processing program; A processor, which, when executing the processing program, implements the clustered network traffic scheduling method based on star flash as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that, The readable storage medium stores a processing program, which, when executed by a processor, implements the clustered network traffic scheduling method based on star flash as described in any one of claims 1 to 7.