Clustering network flow scheduling method and equipment based on star flash, and storage medium

Through the star flash cluster network traffic scheduling method, the model is established using the information of managing G nodes and terminal T nodes, combined with Markov decision-making and multi-agent near-end strategies to train neural networks, optimize resource allocation, and solve the shortcomings of star flash technology in industrial scenarios, achieving high-reliability and low-latency communication effects.

CN120434795APending Publication Date: 2025-08-05SHANGHAI AEROSPACE COMP TECH INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510375076.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing wireless short-range communication technologies such as Bluetooth and WiFi cannot meet the needs of wide coverage, high reliability and low latency in industrial scenarios, and the service flow scheduling strategy of Starflash technology in clustered industrial networks has not been studied in depth.

Method used

The clustered network traffic scheduling method based on star flash is adopted to collect service flow information and cluster node status information by managing G nodes and terminal T nodes, establish a clustered network traffic scheduling model, and combine Markov decision-making and multi-agent near-end strategies to train the neural network model to optimize resource allocation and scheduling decisions.

Benefits of technology

It improves the load balancing capability of clustered network systems, reduces the end-to-end transmission delay of service flows, and ensures high-reliability and low-latency communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434795A_ABST
    Figure CN120434795A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial networks, in particular to a clustering network flow scheduling method based on star flash, which comprises the following steps: S1, collecting service flow information and cluster node state information according to a management G node and a terminal T node; s2, establishing a clustering network flow scheduling model according to the service flow information and the cluster node state information; s3, establishing target optimization and constraint conditions according to the clustering network flow scheduling model; and S4, designing a Markov decision according to target optimization and constraint conditions, training a neural network model according to the Markov decision and a multi-agent near-end strategy, outputting a scheduling decision according to the trained neural network, and obtaining a multi-dimensional resource adaptation result by a solver according to the scheduling decision. And the service flow is transmitted according to the multi-dimensional resource adaptation result and a multi-queue unicast / multicast flow shaping mechanism. According to the invention, the load balancing capability of the star flash clustering network system can be improved, and the end-to-end transmission delay of the service flow can be reduced at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial network technology, and in particular to a clustered network traffic scheduling method, device and storage medium based on Star Flash. Background Art

[0002] In today's era of rapid digitalization and intelligent development, building a smart environment has become a goal pursued in many fields, and wireless short-range communication technology is the core support for achieving this goal. However, traditional wireless short-range communication technologies such as Bluetooth and WiFi suffer from limited coverage, making them unable to meet communication needs over large areas; low accuracy, making it difficult to achieve accurate data transmission; long latency, which makes them poorly perform in scenarios requiring high real-time performance; and poor interference immunity, making communication quality susceptible to external interference. These shortcomings make them difficult to meet the stringent requirements of emerging industrial scenarios for wide coverage, high reliability, and low latency.

[0003] Against this backdrop, Starflash technology emerged. As an emerging wireless short-range communication technology, Starflash introduces a series of new features, such as 5G polar coding, ultra-short frames, and frequency-hopping signal splicing. These new features effectively optimize the transmission range, latency, and anti-interference performance of wireless communication systems, successfully addressing the industry pain points faced by Bluetooth and WiFi, and providing a better solution for the communication needs of emerging industrial scenarios. Despite its significant advantages, Starflash technology only defines the basic architecture and node communication model. When it comes to actual periodic traffic scheduling scenarios in clustered industrial networks, the scheduling strategy for service flows in the Starflash system still requires further in-depth research.

[0004] At the same time, multi-agent deep reinforcement learning technology is gaining popularity in distributed network scenarios. Combining this technology with Starflash scheduling technology promises to achieve more intelligent and dynamic resource allocation. The intelligent decision-making and learning capabilities of multi-agent deep reinforcement learning can better adapt to the complex and ever-changing traffic flows in clustered Starflash networks, thereby improving traffic scheduling performance in clustered Starflash networks and further unlocking the potential of Starflash technology in industrial scenarios. Summary of the Invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and provide a clustered network traffic scheduling method based on Star Flash, which includes the following steps: S1: Collects traffic flow information including traffic flow ID, source address, destination address, packet size, traffic flow period, and latency requirements based on the management G node and terminal T node, as well as cluster node status information; S2: establishing a clustered network traffic scheduling model based on the service flow information and the cluster node status information, wherein the clustered network traffic scheduling model defines a network traffic orchestration period, subframe functions and subframe lengths of different subframes within the orchestration period, and defines a traffic calculation rule on each cluster node according to the orchestration period and subframe function. S3: establishing target optimization and constraints including scheduling period constraints, interval capacity constraints, and traffic transmission constraints according to the clustered network traffic scheduling model; S4: Design a Markov decision based on the target optimization and the constraints, train a neural network model based on the Markov decision and the multi-agent proximal strategy, output a scheduling decision based on the trained neural network, and the solver obtains a multi-dimensional resource adaptation result based on the scheduling decision, and transmits the service flow based on the multi-dimensional resource adaptation result and the multi-queue single multicast traffic shaping mechanism.

[0006] Preferably, in step S1, collecting service flow information and cluster node status information according to the management G node and the terminal T node further includes: Scan and detect network nodes in the network, divide the network into several cluster domains according to the network nodes, the cluster head node in each cluster domain is the management G node of the cluster domain, and the network nodes in the cluster domain other than the management G node are terminal T nodes; The management G node manages and schedules the data transmission of the terminal T nodes in the cluster domain to receive the business flow information uploaded by the terminal T nodes to the management G node of the cluster domain to which they belong, and the management G node manages and collects network status information including link connectivity and bandwidth of the terminal T nodes in the cluster domain in real time.

[0007] Preferably, in step S2, establishing a cluster network traffic scheduling model according to the service flow information and the cluster node status information further includes: The star flash superframe duration is obtained by the sub-time slot scheduling model in the clustered network traffic time slot scheduling model. The sub-time slot scheduling model sets the lowest common multiple of the periodic control traffic cycle as the star flash superframe duration, and the time slot scheduling model requires that the star flash superframe duration is equal to N times the length of the subframe, where N=6, 12, 18, …, 48. The star flash superframe duration is The calculation formula is as follows: The function Used to calculate the least common multiple, Indicates the cycle of each control flow, is a collection of control flows, is the subframe length; The duration of each Star Flash superframe is regarded as a complete traffic scheduling cycle. The Star Flash superframe includes a cluster scheduling subframe for arranging traffic transmission between the terminal T node and the management G node in a cluster domain. and cross-cluster scheduling subframes for orchestrating traffic interactions between G nodes in different clusters ,in, , ,and , Indicates a cluster scheduling subframe across clusters; The cluster scheduling subframe and the cross-cluster scheduling subframe are divided into transmission time slot intervals including uplink transmission interval, downlink transmission interval and protection interval. For the cluster scheduling subframe, if the data communication is to send data from the terminal T node to the management G node, the uplink transmission interval is If the data communication is sending data from the management G node to the terminal T node, it is a downlink transmission interval If the data communication is between the uplink and downlink time slots, it is the protection interval. For the cross-cluster scheduling subframe, since the management G node of one cluster is switched to the terminal T node when the management G node is communicated between two clusters, the transmission interval of the cross-cluster scheduling subframe is divided into an uplink transmission interval, a downlink transmission interval, and a protection interval according to the cluster scheduling subframe principle; When the traffic is scheduled to the mth Star Flash node, the data rate of the node transmitted to the destination Star Flash node d is obtained through the sub-traffic transmission model in the cluster network traffic time slot scheduling model. , the calculation formula is as follows: in, Is an indicator symbol, if the mth node The link is an output link, then , if the mth node The link is an input link, then ,if is a link of non-m nodes, then , It is in the link The data rate of all service flows whose destination address is node d.

[0008] Preferably, in step S3, establishing target optimization according to the clustered network traffic scheduling model further includes: Using the Star Flash superframe duration, the transmission time slot interval, and the data rate, a target optimization for improving the load balancing capability of the Star Flash cluster network system and reducing the end-to-end transmission delay of the service flow is established according to the bandwidth resource utilization and the end-to-end delay of each service flow. The calculation formula is as follows: in, is the bandwidth resource utilization of each link in the uplink and downlink transmission interval, is the end-to-end delay of each service flow, is the weight factor of the sub-goal.

[0009] Preferably, establishing constraint conditions including scheduling period constraint, interval capacity constraint, and traffic transmission constraint according to the clustered network traffic scheduling model further includes: The scheduling period constraint is achieved by planning the traffic transmission within a scheduling period, that is, planning the duration of the star flash superframe; The interval capacity constraint is used to require that the uplink and downlink data transmitted in each subframe interval must not exceed the maximum data rate capacity of the transmission time slot interval. The interval capacity constraint is obtained based on the transmission time slot interval and the service flow information, and is calculated as follows: in, Indicates the size of the service flow data packet. Indicates the link bandwidth, is the transmit power on the link, is the link gain, It is the interference caused by transmitting data through other links. is the length of the transmission interval, Represents a specific transmission interval; The traffic transmission constraint is obtained by calculating that the data rate leaving the network from the destination Star Flash node is equal to the sum of the data rates sent to the destination Star Flash node by all nodes in the network. The calculation formula is as follows: in, is the data rate leaving the destination Star Flash node, and M represents the set of all network nodes that send data streams to network node d.

[0010] Preferably, in step S4, designing a Markov decision according to the target optimization and the constraint conditions further includes: Designing a partially observable state space, an action space, and a cooperation reward for a Markov decision process according to the optimization objective and the constraints; The management G node in each cluster domain observes the service flow information and link bandwidth resource utilization of different terminal T nodes in the respective cluster domain to obtain the observable state space, and the calculation formula is as follows: in, For flow size, For priority, Orchestrate cycles for network traffic, is the source address, For the destination address, is the link bandwidth resource utilization; The action space is obtained based on subframe cluster allocation, transmission time slot interval allocation, bandwidth allocation, and forwarding path allocation. The calculation formula is as follows: Among them, subframe cluster allocation Indicates that a subframe is allocated for communication between nodes in a cluster, that is, only nodes in the specified cluster are allowed to communicate with each other within the subframe duration. Each time slot is uniquely mapped to the uplink transmission interval, the downlink transmission interval and the protection interval, Indicates the amount of bandwidth allocated to users transmitting data within the corresponding interval, Indicates the forwarding path planned for the service flow; The cooperation reward is positively correlated with the load balancing degree and end-to-end delay of the service flow scheduling. The calculation formula of the cooperation reward is as follows: in, , discount factor .

[0011] Preferably, training a neural network model according to the Markov decision and multi-agent proximal strategy further comprises: The multi-agent proximal strategy optimization defines each of the management G nodes as an agent, and the agent maintains a neural network including a strategy estimation neural network for calculating the action corresponding to the current state according to the current strategy and a state evaluation neural network for evaluating the value of the current state. The neural network is initialized by setting the input layer dimension of the strategy estimation neural network to be equal to the dimension of the partially observable state space, the output layer dimension to be equal to the dimension of the action space, and the input layer dimension of the state evaluation neural network to be the same as the dimension of the partially observable state space, and the output layer dimension to be 1. Each management G node observes and collects network state information at each iteration step. The strategy estimation neural network calculates the action and issues the corresponding decision based on the state-action mapping strategy. The network environment executes the action instruction issued by the management G node to obtain the corresponding reward value and transfer to the next state. If the resource utilization of the next state exceeds the capacity of the link time slot or bandwidth, or does not meet the end-to-end delay requirement of the service flow, the scheduling fails and a new round of iteration begins. Otherwise, the current round continues to iterate. The management G node stores the four-tuple of current state, current action, reward, and next state obtained at each step into the experience pool; Among them, after each iteration of state transfer, the management G node batch collects the quadruple in the experience pool to form a trajectory sequence, and at the same time interacts with each other management G node in the environment except itself to obtain the trajectory sequence in the experience pool; Using the trajectory sequence, the target is maximized by the gradient ascent method , to update the strategy to estimate the neural network parameters , maximize the objective The formula is described as: Among them, the parameter T is the length of the trajectory sequence, is the starting step of the trajectory sequence, For the current action, is the current state, is a clipping function that limits the value to In the range, the strategy estimates the old parameters used by the neural network in the previous training step as , the new parameters used in the current step are expressed as , The advantage function value is described as: in, is the discount factor, ; Minimize the loss function by gradient descent method, update the state judgment neural network parameters according to the loss function, the loss function Described as: .

[0012] Based on the same concept, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the star-flash-based clustered network traffic scheduling method as described in any one of the embodiments.

[0013] Based on the same concept, the present invention also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the star-flash-based clustered network traffic scheduling method as described in any one of the embodiments.

[0014] Compared with the prior art, the present invention has the following beneficial effects: The present invention establishes a clustered network traffic scheduling model based on Star Flash through the collected business flow information and cluster node status information, which can effectively integrate network resources and improve resource utilization efficiency. According to the traffic scheduling model, target optimization and constraints including scheduling cycle constraints, interval capacity constraints, and traffic transmission constraints are established to ensure high reliability and low latency communication quality in the clustered network traffic scheduling process.

[0015] The present invention designs Markov decision-making through target optimization and constraint conditions, trains a neural network model based on Markov decision-making and multi-agent proximal strategy, and outputs scheduling decisions based on the trained neural network. This can further improve the scheduling performance of business flows in distributed networks and meet the goals of network load balancing and minimizing end-to-end communication delay. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Various other advantages and benefits will become apparent to those skilled in the art by reading the following detailed description of the preferred embodiment.The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the invention.

[0017] Figure 1 This is a flow chart of the clustered network traffic scheduling method based on Star Flash of the present invention; Figure 2 Another flow chart of the cluster network traffic scheduling method based on Star Flash of the present invention Figure 3 This is an application scenario diagram of the clustered network traffic scheduling method based on Star Flash of the present invention; Figure 4 This is a structural block diagram of the clustered network traffic scheduling device based on Star Flash of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Obviously, the embodiments described are part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative work are within the scope of protection of this application.

[0019] Those skilled in the art will understand that, unless otherwise specified, the singular forms "a," "an," and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0020] First embodiment See also Figure 1 and Figure 2 As shown, this embodiment provides a cluster network traffic scheduling method based on Star Flash, including the following steps: S1: Collects traffic flow information including traffic flow ID, source address, destination address, data packet size, traffic flow period, and delay requirement based on the management G node and terminal T node, as well as cluster node status information.

[0021] Preferably, in step S1, collecting service flow information and cluster node status information according to the management G node and the terminal T node further includes: Scan and detect network nodes in the network, and divide the network into several cluster domains according to the network nodes. The cluster head node in each cluster domain is the management G node of the cluster domain, and the network nodes other than the management G node in the cluster domain are terminal T nodes; The management G node manages and schedules the data transmission of the terminal T nodes in the cluster domain to receive the business flow information uploaded by the terminal T nodes to the management G node of the cluster domain to which they belong. The management G node also manages and collects network status information including link connectivity and bandwidth of the terminal T nodes in the cluster domain in real time.

[0022] See also Figure 3As shown, the network topology provided in this embodiment is divided into four cluster domains. Each cluster includes a cluster head management node (G) and cluster member nodes (T). Nodes communicate with each other over short distances within a range of 10 to 20 meters. The transmit power is set to 10 dBm, and the channel frequency band uses a 5 GHz band to support high-speed data transmission, with a frequency bandwidth of 20 MHz. Clusters 1 and 2 are separated by a barrier, preventing direct communication between the two clusters. Communication between the two clusters requires forwarding through multi-hop relay nodes. Each T node periodically transmits control traffic flows, with data packets ranging in size from 50 bytes to 1 KB. To optimize network traffic scheduling performance, the time slot intervals, paths, and bandwidth required for data forwarding are determined by a traffic scheduling algorithm coordinated by multiple management nodes. This improves the system's load balancing capabilities and optimizes the end-to-end latency of traffic flow transmission. The Adam optimizer is used with a learning rate of 1e-4. The middle layer consists of two fully connected neural networks. The experience pool size is set to 1024, the sampling batch size is set to 32, and the discount factor is set to 0.95.

[0023] S2: A clustered network traffic scheduling model based on Star Flash is established based on the business flow information and cluster node status information. The clustered network traffic scheduling model defines the network traffic orchestration period, the subframe functions and subframe lengths of different subframes within the orchestration period, and defines the calculation rules for the traffic on each cluster node according to the orchestration period and subframe function.

[0024] Preferably, in step S2, establishing a clustered network traffic scheduling model based on the service flow information and cluster node status information further includes: The duration of the star flash superframe is obtained by the sub-time slot scheduling model in the clustered network traffic time slot scheduling model. The sub-time slot scheduling model sets the lowest common multiple of the periodic control traffic cycle as the star flash superframe duration, and the time slot scheduling model requires that the star flash superframe duration is equal to N times the length of the subframe, where N=6, 12, 18, …, 48. The star flash superframe duration The calculation formula is as follows: The function Used to calculate the least common multiple, Indicates the cycle of each control flow, is a collection of control flows, is the subframe length; The duration of each Star Flash superframe is regarded as a complete traffic scheduling cycle. The Star Flash superframe includes cluster scheduling subframes used to arrange traffic transmission between terminal T nodes and management G nodes within a cluster domain. and cross-cluster scheduling subframes for orchestrating traffic interactions between G nodes in different clusters ,in, , ,and , Indicates a cluster scheduling subframe across clusters; The transmission time slot intervals including uplink transmission interval, downlink transmission interval and protection interval are divided for cluster scheduling subframes and cross-cluster scheduling subframes. Among them, for cluster scheduling subframes, if the data communication is sending data from the terminal T node to the management G node, it is the uplink transmission interval If the data communication is from the management G node to the terminal T node, it is the downlink transmission interval. If the data communication is between the uplink and downlink time slots, it is the protection interval. , where the guard interval between the uplink and downlink time slots is used to avoid conflicts between the uplink and downlink time slots; for cross-cluster scheduling subframes, since the management G node of one cluster is switched to the terminal T node when the two clusters communicate with each other, the transmission interval of the cross-cluster scheduling subframe is divided into uplink transmission interval, downlink transmission interval and guard interval according to the cluster scheduling subframe principle; When the traffic is scheduled to the mth Star Flash node, the data rate of the node transmitted to the destination Star Flash node d is obtained through the sub-traffic transmission model in the clustered network traffic time slot scheduling model. , the calculation formula is as follows: in, Is an indicator symbol, if the mth node The link is an output link, then , if the mth node The link is an input link, then ,if is a link of non-m nodes, then , It is in the link The data rate of all service flows whose destination address is node d.

[0025] S3: Establish target optimization and constraints including scheduling period constraints, interval capacity constraints, and traffic transmission constraints based on the clustered network traffic scheduling model.

[0026] Preferably, in step S3, establishing target optimization according to the clustered network traffic scheduling model further includes: Using the Star Flash superframe duration, transmission time slot interval, and data rate, we establish a target optimization for improving the load balancing capability of the Star Flash cluster network system and reducing the end-to-end transmission delay of the service flow based on bandwidth resource utilization and the end-to-end delay of each service flow. The calculation formula is as follows: in, is the bandwidth resource utilization of each link in the uplink and downlink transmission interval, is the end-to-end delay of each service flow, is the weight factor of the sub-goal.

[0027] Preferably, establishing constraints including scheduling period constraints, interval capacity constraints, and traffic transmission constraints according to the clustered network traffic scheduling model further includes: Scheduling period constraints are implemented by planning traffic transmission within an orchestration period (also known as a scheduling period), that is, planning the duration of the Star Flash superframe; The interval capacity constraint is used to require that the uplink and downlink data transmitted in each subframe interval must not exceed the maximum data rate capacity of the transmission time slot interval. The interval capacity constraint is obtained based on the transmission time slot interval and service flow information. The calculation formula is as follows: in, Indicates the size of the service flow data packet. Indicates the link bandwidth, is the transmit power on the link, is the link gain, It is the interference caused by transmitting data through other links. is the length of the transmission interval, Represents a specific transmission interval; The traffic transmission constraint is obtained by calculating the data rate leaving the network from the destination Star Flash node equal to the sum of the data rates sent to the destination Star Flash node by all nodes in the network. The calculation formula is as follows: in, is the data rate leaving the destination Star Flash node, and M represents the set of all network nodes that send data streams to network node d.

[0028] S4: Design a Markov decision based on the target optimization and constraints, train a neural network model based on the Markov decision and multi-agent proximal strategy, output a scheduling decision based on the trained neural network, and the solver obtains a multi-dimensional resource adaptation result based on the scheduling decision. The service flow is transmitted based on the multi-dimensional resource adaptation result and the multi-queue single multicast traffic shaping mechanism.

[0029] Preferably, in step S4, designing a Markov decision according to the target optimization and the constraint conditions further includes: According to the target optimization and constraints, the partially observable state space, action space and cooperation reward of the Markov decision process are designed. Specifically, in this embodiment, based on the target optimization and constraints of the optimization problem, the partially observable state space, action space and cooperation reward of the Markov decision process are designed. Each cluster head G node observes the traffic characteristics of different T nodes in its own cluster, including traffic size, cycle, address information, priority, etc. The management G node in each cluster domain observes the service flow information and link bandwidth resource utilization of different terminal T nodes in its own cluster domain to obtain the observable state space. The calculation formula is as follows: in, For flow size, For priority, Orchestrate cycles for network traffic, is the source address, For the destination address, is the link bandwidth resource utilization; The action space is obtained based on subframe cluster allocation, transmission time slot interval allocation, bandwidth allocation, and forwarding path allocation. The calculation formula is as follows: Among them, subframe cluster allocation Indicates that a subframe is allocated for communication between nodes in a cluster, that is, only nodes in the specified cluster are allowed to communicate with each other within the subframe duration. Each time slot is uniquely mapped to the uplink transmission interval, downlink transmission interval and protection interval, Indicates the amount of bandwidth allocated to users transmitting data within the corresponding interval, Indicates the forwarding path planned for the service flow; The cooperation reward is positively correlated with the load balancing degree and end-to-end delay of the service flow scheduling. The calculation formula for the cooperation reward is as follows: in, , discount factor .

[0030] Preferably, training the neural network model according to Markov decision making and multi-agent proximal strategy further comprises: Multi-agent proximal strategy optimization defines each management G node as an agent. The agent maintains a neural network including a policy estimation neural network for calculating the action corresponding to the current state according to the current policy and a state evaluation neural network for evaluating the value of the current state. The neural network is initialized by setting the input layer dimension of the policy estimation neural network equal to the dimension of the partially observable state space, the output layer dimension equal to the dimension of the action space, and the input layer dimension of the state evaluation neural network equal to the dimension of the partially observable state space, and the output layer dimension to 1. Specifically, in this embodiment, the intermediate layers all include K layers of fully connected neural networks. Each management G node observes and collects network state information at each iteration step. The policy estimation neural network calculates actions and issues corresponding decisions based on the state-action mapping strategy. The network environment executes the action instructions issued by the management G node, obtains the corresponding reward value, and moves to the next state. If the resource utilization of the next state exceeds the link time slot or bandwidth capacity, or does not meet the end-to-end delay requirements of the service flow, the scheduling fails and a new round of iteration begins. Otherwise, the current round continues to iterate. The management G node stores the four-tuple of current state, current action, reward, and next state obtained at each step in the experience pool. Among them, after each iteration of state transfer, the management G node batch collects the four-tuples in the experience pool to form a trajectory sequence, and at the same time interacts with each other management G node in the environment except itself to obtain the trajectory sequence in the experience pool; Using the trajectory sequence, the target is maximized by the gradient ascent method , to update the strategy to estimate the neural network parameters , maximize the objective The formula is described as: Among them, the parameter T is the length of the trajectory sequence, is the starting step of the trajectory sequence, For the current action, is the current state, is a clipping function that limits the value to In the range, the old parameters used by the policy estimation neural network in the previous training step are expressed as , the new parameters used in the current step are expressed as , The advantage function value is described as: in, is the discount factor, ; Minimize the loss function by gradient descent method, update the state judgment neural network parameters according to the loss function, the loss function Described as: .

[0031] The embodiment of the present invention also provides a cluster network traffic scheduling device based on star flash, such as Figure 4 As shown, the device includes: Cluster node status and service flow information collection module, used to collect service flow information including service flow ID, source address, destination address, data packet size, service flow data volume, priority, service flow cycle and delay requirement of terminal T nodes in the cluster through management G nodes, as well as cluster node status information including connection status, time slot and bandwidth utilization; The module for building a traffic scheduling model for clustered networks based on Starflash is used to construct the frame structure and traffic transmission model for traffic scheduling in the Starflash clustered network system. Based on the frame structure and traffic transmission model, the module obtains the Starflash superframe duration, transmission time slot interval, and data rate. An optimization problem construction module based on the Star Flash clustering network traffic scheduling model is used to establish target optimization and constraints including scheduling period constraints, interval capacity constraints, and traffic transmission constraints based on the Star Flash superframe duration, transmission time slot interval, and data rate; A Markov decision process building module for multi-management node collaboration, which is used to establish the state space, action space, and reward of the Markov decision process based on target optimization and constraints, providing theoretical support for the design module of the traffic scheduling algorithm for multi-management node collaboration; A traffic scheduling algorithm design module for multi-management node collaboration is used to train a neural network model based on Markov decision making and multi-agent proximal strategies. The trained neural network outputs scheduling decisions to achieve periodic control of traffic within and between scheduling clusters, thereby improving system scheduling performance. The solution module is used to obtain a multi-dimensional resource adaptation result according to the scheduling decision, and transmit the service flow according to the multi-dimensional resource adaptation result and the multi-queue single multicast traffic shaping mechanism.

[0032] Second embodiment In some embodiments of the present application, a computer device is also provided, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor performs the steps of the star-flash-based clustered network traffic scheduling method as described in any one of the embodiments.

[0033] The present invention also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, enables one or more processors to perform the steps of the Starflash-based clustered network traffic scheduling method as described in any one of Example 1.

[0034] Based on the same concept, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the star-flash-based clustered network traffic scheduling method as described in any one of the embodiments.

[0035] It can be understood that for the aforementioned star-flash-based clustered network traffic scheduling method, if it is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer server, or a network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0036] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0037] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A cluster network traffic scheduling method based on star flash, characterized in that: The following steps are involved: S1: Collects traffic flow information including traffic flow ID, source address, destination address, packet size, traffic flow period, and latency requirements based on the management G node and terminal T node, as well as cluster node status information; S2: establishing a clustered network traffic scheduling model based on the service flow information and the cluster node status information, wherein the clustered network traffic scheduling model defines a network traffic orchestration period, subframe functions and subframe lengths of different subframes within the orchestration period, and defines a traffic calculation rule on each cluster node according to the orchestration period and subframe function. S3: establishing target optimization and constraints including scheduling period constraints, interval capacity constraints, and traffic transmission constraints according to the clustered network traffic scheduling model; S4: Design a Markov decision based on the target optimization and the constraints, train a neural network model based on the Markov decision and the multi-agent proximal strategy, output a scheduling decision based on the trained neural network, and the solver obtains a multi-dimensional resource adaptation result based on the scheduling decision, and transmits the service flow based on the multi-dimensional resource adaptation result and the multi-queue single multicast traffic shaping mechanism.

2. The cluster network traffic scheduling method based on star flash according to claim 1 is characterized in that: In step S1, the service flow information and cluster node status information are collected according to the management G node and the terminal T node, further comprising: Scan and detect network nodes in the network, divide the network into several cluster domains according to the network nodes, the cluster head node in each cluster domain is the management G node of the cluster domain, and the network nodes in the cluster domain other than the management G node are terminal T nodes; The management G node manages and schedules the data transmission of the terminal T nodes in the cluster domain to receive the business flow information uploaded by the terminal T nodes to the management G node of the cluster domain to which they belong, and the management G node manages and collects network status information including link connectivity and bandwidth of the terminal T nodes in the cluster domain in real time.

3. The cluster network traffic scheduling method based on star flash according to claim 2 is characterized in that: In step S2, a cluster network traffic scheduling model is established according to the service flow information and the cluster node status information, further comprising: The star flash superframe duration is obtained by the sub-time slot scheduling model in the clustered network traffic scheduling model. The sub-time slot scheduling model sets the lowest common multiple of the periodic control traffic cycle as the star flash superframe duration, and the time slot scheduling model requires that the star flash superframe duration is equal to N times the length of the subframe, where N=6, 12, 18, …, 48. The star flash superframe duration is The calculation formula is as follows: The function Used to calculate the least common multiple, Indicates the cycle of each control flow, is a collection of control flows, is the subframe length; The duration of each Star Flash superframe is regarded as a complete network traffic orchestration cycle. The Star Flash superframe includes a cluster scheduling subframe for orchestrating traffic transmission between the terminal T node and the management G node in a cluster domain. and cross-cluster scheduling subframes for orchestrating traffic interactions between G nodes in different clusters ,in, , ,and , Indicates a cluster scheduling subframe across clusters; The cluster scheduling subframe and the cross-cluster scheduling subframe are divided into transmission time slot intervals including uplink transmission interval, downlink transmission interval and protection interval. For the cluster scheduling subframe, if the data communication is to send data from the terminal T node to the management G node, the uplink transmission interval is If the data communication is sending data from the management G node to the terminal T node, it is a downlink transmission interval If the data communication is between the uplink and downlink time slots, it is the protection interval. For the cross-cluster scheduling subframe, since the management G node of one cluster is switched to the terminal T node when the management G node is communicated between two clusters, the transmission interval of the cross-cluster scheduling subframe is divided into an uplink transmission interval, a downlink transmission interval, and a protection interval according to the cluster scheduling subframe principle; When the traffic is scheduled to the mth Star Flash node, the data rate of the node transmitted to the destination Star Flash node d is obtained through the sub-traffic transmission model in the cluster network traffic time slot scheduling model. , the calculation formula is as follows: in, Is an indicator symbol, if the mth node The link is an output link, then , if the mth node The link is an input link, then ,if is a link of non-m nodes, then , It is in the link The data rate of all service flows whose destination address is node d.

4. The cluster network traffic scheduling method based on star flash according to claim 3 is characterized in that: In step S3, establishing target optimization according to the clustered network traffic scheduling model further includes: Using the Star Flash superframe duration, the transmission time slot interval, and the data rate, a target optimization for improving the load balancing capability of the Star Flash cluster network system and reducing the end-to-end transmission delay of the service flow is established according to the bandwidth resource utilization and the end-to-end delay of each service flow. The calculation formula is as follows: in, is the bandwidth resource utilization of each link in the uplink and downlink transmission interval, is the end-to-end delay of each service flow, is the weight factor of the sub-goal.

5. The cluster network traffic scheduling method based on star flash according to claim 4 is characterized in that: Establishing constraints including scheduling period constraints, interval capacity constraints, and traffic transmission constraints according to the clustered network traffic scheduling model further includes: The scheduling period constraint is achieved by planning the traffic transmission within an orchestration period, that is, planning the duration of the star flash superframe; The interval capacity constraint is used to require that the uplink and downlink data transmitted in each subframe interval must not exceed the maximum data rate capacity of the transmission time slot interval. The interval capacity constraint is obtained based on the transmission time slot interval and the service flow information, and is calculated as follows: in, Indicates the size of the service flow data packet. Indicates the link bandwidth, is the transmit power on the link, is the link gain, It is the interference caused by transmitting data through other links. is the length of the transmission interval, Represents a specific transmission interval; The traffic transmission constraint is obtained by calculating that the data rate leaving the network from the destination Star Flash node is equal to the sum of the data rates sent to the destination Star Flash node by all nodes in the network. The calculation formula is as follows: in, is the data rate leaving the destination Star Flash node, and M represents the set of all network nodes that send data streams to network node d.

6. The cluster network traffic scheduling method based on star flash according to claim 5 is characterized in that: In step S4, designing a Markov decision according to the target optimization and the constraint conditions further includes: Designing a partially observable state space, an action space, and a cooperation reward for a Markov decision process according to the optimization objective and the constraints; The management G node in each cluster domain observes the service flow information and link bandwidth resource utilization of different terminal T nodes in the respective cluster domain to obtain the observable state space, and the calculation formula is as follows: in, For flow size, For priority, Orchestrate cycles for network traffic, is the source address, For the destination address, is the link bandwidth resource utilization; The action space is obtained based on subframe cluster allocation, transmission time slot interval allocation, bandwidth allocation, and forwarding path allocation. The calculation formula is as follows: Among them, subframe cluster allocation Indicates that a subframe is allocated for communication between nodes in a cluster, that is, only nodes in the specified cluster are allowed to communicate with each other within the subframe duration. Each time slot is uniquely mapped to the uplink transmission interval, the downlink transmission interval and the protection interval, Indicates the amount of bandwidth allocated to users transmitting data within the corresponding interval, Indicates the forwarding path planned for the service flow; The cooperation reward is positively correlated with the load balancing degree and end-to-end delay of the service flow scheduling. The calculation formula of the cooperation reward is as follows: in, , discount factor .

7. The cluster network traffic scheduling method based on star flash according to claim 6 is characterized in that: Training the neural network model according to the Markov decision and multi-agent proximal strategy further includes: The multi-agent proximal strategy optimization defines each of the management G nodes as an agent, and the agent maintains a neural network including a strategy estimation neural network for calculating the action corresponding to the current state according to the current strategy and a state evaluation neural network for evaluating the value of the current state. The neural network is initialized by setting the input layer dimension of the strategy estimation neural network to be equal to the dimension of the partially observable state space, the output layer dimension to be equal to the dimension of the action space, and the input layer dimension of the state evaluation neural network to be the same as the dimension of the partially observable state space, and the output layer dimension to be 1. Each management G node observes and collects network state information at each iteration step. The strategy estimation neural network calculates the action and issues the corresponding decision based on the state-action mapping strategy. The network environment executes the action instruction issued by the management G node to obtain the corresponding reward value and transfer to the next state. If the resource utilization of the next state exceeds the capacity of the link time slot or bandwidth, or does not meet the end-to-end delay requirement of the service flow, the scheduling fails and a new round of iteration begins. Otherwise, the current round continues to iterate. The management G node stores the four-tuple of current state, current action, reward, and next state obtained at each step into the experience pool; Among them, after each iteration of state transfer, the management G node batch collects the quadruple in the experience pool to form a trajectory sequence, and at the same time interacts with each other management G node in the environment except itself to obtain the trajectory sequence in the experience pool; Using the trajectory sequence, the gradient ascent method is used to maximize the target , to update the strategy to estimate the neural network parameters , maximize the objective The formula is described as: Among them, the parameter T is the length of the trajectory sequence, is the starting step of the trajectory sequence, For the current action, is the current state, is a clipping function that limits the value to In the range, the strategy estimates the old parameters used by the neural network in the previous training step as , the new parameters used in the current step are expressed as , The advantage function value is described as: in, is the discount factor, ; Minimize the loss function by gradient descent method, update the state judgment neural network parameters according to the loss function, the loss function Described as: 。 8. A computer device, characterized in that: It includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the star-flash-based clustered network traffic scheduling method as described in any one of claims 1 to 7.

9. A storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to perform the steps of the clustered network traffic scheduling method based on Starflash as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Link distribution method, device, equipment and product

    CN121001146A