Multi-agent cooperative scheduling method for incremental mixed flow in industrial time-sensitive network

By employing a multi-agent collaborative scheduling method, the dynamic interference problem caused by the incremental mixing of streams in industrial time-sensitive networks was solved, transmission assurance capabilities were improved, the latency and bandwidth requirements of audio and video bridging traffic in the network were ensured, and efficient mixing stream management was achieved.

CN121396918AActive Publication Date: 2026-01-23NORTHEASTERN UNIV CHINA +1

Patent Information

Application Number
CN202511902330.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-01-23
Estimated Expiration
2045-12-17

AI Technical Summary

Technical Problem

In industrial time-sensitive networks, the incremental introduction of mixed streams causes low-priority audio and video bridging (AVB) traffic to fail to meet latency constraints, and the transmission bandwidth is squeezed, affecting network stability and transmission efficiency.

Method used

A multi-agent cooperative scheduling method is adopted. By finely modeling the dynamic interference caused by the hybrid flow increment, a traffic scheduling model based on multi-agent cooperation is designed to generate an end-to-end transmission scheme that meets the time and bandwidth constraints. The transmission decision is optimized by using a traffic-aware neural network (FANet).

Benefits of technology

It enhances the transmission guarantee capability of industrial time-sensitive networks for incremental mixed streams, ensuring the data transmission needs of heterogeneous industrial applications and achieving efficient scheduling and management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121396918A_ABST
    Figure CN121396918A_ABST
Patent Text Reader

Abstract

The multi-agent cooperative scheduling method for the incremental mixed flow in the industrial time-sensitive network comprises the following steps: analyzing transmission characteristics of an AVB flow and a TT flow in the industrial time-sensitive network, and carrying out refined modeling on dynamic interference caused by the increment of the mixed flow based on a network calculation theory; determining key transmission management variables of the AVB stream and the TT stream in combination with a dynamic interference factor introduced by the incremental mixed stream; designing a flow scheduling model based on multi-agent cooperation, and generating an end-to-end transmission scheme based on multi-agent cooperation decision when incremental mixed flow is accessed; and packaging the trained agents into a multi-agent cluster and deploying the multi-agent cluster into an industrial time-sensitive network for generating a transmission scheme for the incremental mixed flow on line. According to the method, the transmission requirement of the incremental mixed flow can be met while the existing AVB flow delay constraint is guaranteed, the transmission scheme is optimized to improve the service quality of an industrial time-sensitive network, and stable operation of advanced applications in the industrial Internet of Things is fully supported.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of industrial Internet of Things, and relates to a multi-agent cooperative scheduling method for incremental mixed flow in an industrial time-sensitive network. BACKGROUND

[0002] In recent years, Time-Sensitive Networking (TSN) focusing on low latency and high reliability communication has been extended to the wireless domain, building an industrial Internet of Things system with flexible access capability and high reliability. In the industrial Internet of Things enabled by TSN, various advanced industrial applications such as robot collaboration, remote control, and intelligent human-computer interaction are booming, relying on the deterministic transmission service provided by TSN. These advanced industrial applications are often multifunctional, and the various differentiated industrial data required by them are transmitted through TSN traffic of different priorities, resulting in a large number of mixed flows in the network. In addition, the industrial Internet of Things supports the dynamic deployment of customized applications in the industrial field without interrupting operation to extend industrial functions in real time, which means that mixed flows in the industrial Internet of Things enabled by TSN will continue to increase. However, in the TSN standard, the priority of Audio-Video Bridging (AVB) traffic is lower than that of Time-Triggered (TT) traffic, and the available bandwidth of AVB traffic will be squeezed due to the increase of mixed flows, thereby increasing the end-to-end latency. To ensure the stable operation of industrial applications in the network, the increase of mixed flows should not cause the transmission failure of existing AVB flows due to the failure to meet the latency constraint; on the other hand, whether the incremental flow can be successfully transmitted also depends on whether the transmission scheme provided for it meets its own transmission requirements. Therefore, there is an urgent need for a mixed flow transmission management strategy that takes into account the above two conditions to optimize the transmission scheme of incremental mixed flows. SUMMARY

[0003] To solve the above technical problems, the purpose of the present application is to provide a multi-agent cooperative scheduling method for incremental mixed flow in an industrial time-sensitive network, which first provides perfect modeling of the dynamic interference introduced by the increase of mixed flow, and then designs a traffic scheduling mechanism based on multi-agent cooperation, so as to improve the guarantee capability of the industrial time-sensitive network for the transmission of incremental mixed flow and meet the data transmission requirements of heterogeneous industrial applications.

[0004] The present application provides a multi-agent cooperative scheduling method for incremental mixed flow in an industrial time-sensitive network, comprising:

[0005] Step 1: analyze the transmission characteristics of AVB flows and TT flows in the industrial time-sensitive network, and based on the theory of network calculus, combine the bandwidth competition and latency fluctuation under the scenario of incremental mixed flow to finely model the dynamic interference caused by the increase of mixed flow;

[0006] Step 2: Combine the dynamic interference factors introduced by the incremental mixed flow, and determine the key transmission management variables of the AVB flow and the TT flow, covering transmission route determination and scheduling time slot division;

[0007] Step 3: Design a traffic scheduling model based on multi-agent collaboration, and generate an end-to-end transmission scheme that meets the delay and bandwidth constraints based on multi-agent collaborative decision-making when the incremental mixed flow is accessed;

[0008] Step 4: Encapsulate the trained agent as a multi-agent cluster and deploy it to the industrial time-sensitive network for online generation of transmission schemes for incremental mixed flows.

[0009] The multi-agent collaborative scheduling method for incremental mixed flows in an industrial time-sensitive network of the present application first analyzes the transmission mode of mixed traffic in the industrial time-sensitive network, and on this basis, models the dynamic interference caused by incremental mixed transmission and determines the mixed flow transmission management variables; then designs a traffic scheduling method based on multi-agent collaboration, and configures a traffic adaptive neural network for the agent. This method can improve the transmission guarantee capability of the industrial time-sensitive network for incremental mixed flows, meet the data transmission needs of heterogeneous industrial applications, and achieve efficient scheduling and management of incremental mixed flows. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 is a flowchart of a multi-agent collaborative scheduling method for incremental mixed flows in an industrial time-sensitive network of the present application. DETAILED DESCRIPTION

[0011] As Figure 1 shown, the multi-agent collaborative scheduling method for incremental mixed flows in an industrial time-sensitive network of the present application comprises:

[0012] Step 1: Analyze the transmission characteristics of AVB flows and TT flows in the industrial time-sensitive network, including the priority mechanism and bandwidth occupation mode of the two; based on network calculus theory, combined with bandwidth competition and delay fluctuation in the incremental mixed flow scenario, model the dynamic interference caused by the incremental mixed flow in detail.

[0013] In the industrial time-sensitive network, the determinism of the TT flow hop-by-hop transmission is based on the reservation of dedicated time slots for it on each link, while the AVB flow is provided with a lower level of determinism due to its lower priority and is not allocated a time slot. This means that once the TT flow is allocated a time slot, it will not be disturbed by the bandwidth occupation that may be caused by subsequent incremental mixed flows. However, this is not the case for AVB. Whether it is an incremental TT flow or an incremental AVB flow, it will compress the available bandwidth of the existing AVB flow, thereby affecting its end-to-end delay and transmission. This phenomenon is a direct manifestation of the dynamic interference introduced by the incremental mixed flow.

[0014] The topology of a TSN network is denoted as is a set of network nodes, is a set of directed links:

[0015]

[0016] wherein, denote the start node, end node and bandwidth of link is a set of ports in the network, denotes the source port corresponding to link For any port , the links with this port as source are denoted as

[0017] In the advanced intelligent manufacturing scenario, the on-demand deployment of industrial applications will incrementally introduce mixed flows containing AVB flows and TT flows into the TSN network, causing denotes the index sequence of the incremental mixed flows, and the th incremental mixed flow is denoted as , which is defined as follows:

[0018]

[0019] wherein, denotes the start point of the incremental mixed flow, denotes the end point of the incremental mixed flow, denotes the data size of the incremental mixed flow, denotes the transmission period of the incremental mixed flow, denotes the type identifier of the incremental mixed flow, denotes the upper limit of the tolerable end-to-end delay of the incremental mixed flow.

[0020] When AVB flows or TT flows are incrementally deployed in the network, the transmission process of the existing AVB flows in the network will be disturbed, and their transmission performance will also be affected. Specifically, if the new flow is an AVB flow, it will exacerbate the bandwidth contention among all network AVB flows; and if the new flow is a TT flow, its occupation of a specific time slot will compress the available bandwidth of the AVB flow. The addition of mixed flows will increase the single-hop delay of the existing AVB flows; and then it may cause the upper limit of the end-to-end delay to be broken, resulting in unacceptable transmission performance. Based on this analysis, by modeling the delay in the single-hop transmission process of AVB flows, the dynamic disturbance caused by the incremental mixed flows can be effectively described.

[0021] ​​​​​According to network calculus theory, the upper limit of the single-hop delay of an AVB stream transmitted through a port is the maximum difference in the horizontal direction between the service curve and the arrival curve of the AVB stream transmitted through a certain port.

[0022] port Service curves for AVB streaming for:

[0023]

[0024] in, Indicates link bandwidth; Indicates port Time offset relative to network startup time At this location, the cumulative active duration of the dedicated timeslot and guard band for the TT stream is displayed. This function reflects the unavailable time window of the AVB stream and can be accessed by resolving the port. The gate control list is obtained directly. During the period when the AVB stream is not transmissible, The function value increases linearly with time; however, it remains constant during the period when the AVB stream can be transmitted.

[0025] For the network index is AVB stream Assume that it is generated by the terminal and passes through the first hop port. Upon joining the network, it is in The arrival curve can be characterized by the bucket-and-slot model:

[0026]

[0027] in, express The average generation rate within its cycle, Indicates its maximum possible instantaneous transmission volume; for Any port traversed in subsequent hops , its in The arrival curve is the previous hop port. Output curve:

[0028]

[0029] in, For flow In the port and its corresponding links The upper limit of single-hop delay during transmission; this value is added to the time synchronization error between nodes. Afterwards, as bias embedding exist arrival curve To depict the port Flow Output to port The effect.

[0030] When the incremental flow According to the transmission scheme specified for it When deployed on a network, it will cause updates to the service curves or arrival curves of some ports in the network: if For AVB flows, incremental deployment updates the aggregated AVB flows and their arrival curves on some ports, thereby increasing the single-hop latency of existing AVB flows in the network; if For TT streams, incremental deployment may compress the AVB stream transmission window and update the port service curve, thereby changing the upper bound of the single-hop latency of existing AVB streams.

[0031] Step 2: Considering the dynamic interference factors introduced by incremental hybrid streams, identify the key transmission management variables for AVB and TT streams, covering transmission route determination and scheduling time slot allocation, specifically:

[0032] Step 2.1: The difference between AVB and TT streams in whether they are allocated dedicated time slots leads to different transmission schemes: AVB streams only need to specify the route, while TT streams must determine the joint route and time slot scheduling scheme hop-by-hop. (Note the hybrid stream increment.) The transmission scheme is Its definition is:

[0033]

[0034] in, For incremental mixed flow The number of transmission hops, for In other words, Indicates incremental mixed flow In the The links traversed during the jump, For link set The Kleene positive closure represents the set of all potential routes with a hop count of at least one; when When it is a TT stream, considering that it needs to be transmitted within one supercycle... Then, it must be allocated on each link it passes through. One time slot; Indicates in Up to be assigned to TT stream of The index sequence of time slots, family It contains all The possible sequences of time slot indices, and denotes with the Cartesian product of the Kleene positive closure of ; each item in this closure corresponds to a potential end-to-end transmission scheme of TT flows with a joint routing-scheduling combination.

[0035] Step 2.2: To ensure the link order connection in traffic routing and avoid forming a loop, the routing in the transmission scheme needs to satisfy the following constraints:

[0036]

[0037] wherein, denotes the start node of link , denotes the end node of link , the end node of link , denotes the start node of link .

[0038] When is a TT flow, the time slots allocated to it on each link need to further satisfy the following constraints:

[0039]

[0040] wherein, is the th index in , corresponding to the time slot scheduled for the flow for the th transmission within a super cycle; H denotes the super cycle configuration of the TSN network, denotes the duration of a time slot; the offset between each pair of adjacent time slots must strictly match the flow period to ensure the periodicity of TT flow transmission.

[0041] Step 3: Design a traffic scheduling model based on multi-agent collaboration. When incremental mixed flows are accessed, generate an end-to-end transmission scheme that satisfies the delay and bandwidth constraints based on multi-agent collaborative decision-making. Specifically:

[0042] Step 3.1: Construct a multi-agent Markov decision process.

[0043] Convert the solution of the incremental mixed flow transmission management problem into a multi-agent Markov decision process. In this multi-agent Markov decision process, let denote the state space. When making transmission decisions for a certain incremental mixed flow, This represents the initial state quantity, which includes flow attribute information and the overall state of the network. The action space is shared by all agents; the incremental hybrid flow is formed through hop-by-hop cooperation among agents. In constructing an end-to-end transmission scheme, the output action of each agent involves the routing selection of the current hop, thus specifying which adjacent agent will continue the decision for the next hop. If the agents participate in sequence and cooperate to complete the shared... The jump-by-jump decision-making process is represented as follows:

[0044]

[0045] in, Indicates the decision-making process. To characterize the overall agent action of the end-to-end transmission scheme, corresponding to incremental hybrid streams end-to-end transmission scheme , For the first A smart agent The output single-hop transmission action, for the agent In other words, Defined as from the previous The action chain formed by the sequential decision-making outputs of each agent guides and constrains its current single-hop transmission decision; the overall decision of multiple agents can be represented as an action sequence composed of the single-hop actions of each agent, which corresponds to the end-to-end transmission scheme.

[0046] Step 3.2: Design the state input for the multi-agent system.

[0047] For the An intelligent agent participating in decision-making Its state input is determined by the initial state. With action chain Composition; initial state Contains Stream Attribute information (e.g., flow type, period, latency constraints, bandwidth requirements, etc.), and overall network information organized based on topological connectivity. (For example, the current transmission load of each link, time slot utilization, etc.); for intelligent agents Action chain It is concretized into the local observation state in the local network segment it maps to. This local observation characterizes the cumulative effect of preceding decisions on adjacent links.

[0048] In local observations, if incremental mixing flow For AVB flows, the agent focuses on the one-hop delay upper bound of both AVB flows on the adjacent link, and evaluates the potential impact on the link load via the transmission on the link; if For TT flows, the agent refines the observation granularity to the available time slots on the adjacent link, analyzes the fluctuations in the link transmission load and time slot utilization under different time slot allocation choices.

[0049] Initial state Local state Complete input formed by combination As the state input of the agent in hop-by-hop cooperation.

[0050] Step 3.3: Define differentiated action output modes for the agent for different incremental flow types.

[0051] For AVB flows, the agent outputs actions under the state The action For AVB flows, the agent outputs actions under the state The action contains both the next-hop transmission link and the set of time slots allocated on the link to ensure that the routing scheme and time slot allocation scheme of TT flows are compatible and conflict-free under the joint decision framework of routing and time slot scheduling. Through the above design, the agent is provided with a differentiated action space description and definition for incremental AVB flows and TT flows.

[0052] Step 3.4: To enable multiple agents to jointly optimize the end-to-end transmission performance of incremental mixed flows in the cooperation process, construct a reward function based on task achievement.

[0053] When multiple agents successfully generate an end-to-end transmission scheme that meets the bandwidth and delay constraints for a certain incremental flow through cooperation, the transmission scheme is considered as a valid configuration, and a positive reward is given to all agents involved in the decision-making process in the hop-by-hop decision-making process, with the reward value set to +1; if any hop in the decision-making chain cannot find a feasible next-hop link or available time slot, resulting in the inability to form a complete end-to-end transmission scheme, the cooperation is considered a failure, and a negative reward is given to all participating agents, with the reward value set to -1.

[0054] Step 3.5: Configure a flow-aware neural network (FANet) for each agent to realize the policy mapping from state input to action output, specifically:

[0055] For state input initial state Flow characteristics and network characteristics The system consists of two parts: a shared traffic feature extraction layer and a network feature extraction layer to extract high-dimensional features related to flow attributes and the overall network state; and a heterogeneous feature extraction layer to adaptively handle the local states corresponding to different traffic types. This allows for handling different dimensions under different conditions, such as AVB streams and TT streams. Feature extraction is performed; after feature extraction, the hidden features output from each feature extraction layer are concatenated, and the concatenated features are input into the corresponding decision module according to the traffic type, outputting the final action distribution or action selection; the agent... The decision-making strategy is represented as ,in for The policy parameters of a neural network.

[0056] Step 3.6: Train a multi-agent policy based on the proximal policy optimization algorithm; integrate the policy parameters of all agents into... Based on this, the overall decision-making strategy of the intelligent agent cluster is defined as follows: In the multi-agent Markov decision-making process, the agents obtain cumulative rewards by generating transmission schemes for incremental hybrid streams. And by leveraging the advantage function To measure in the initial state Take action below Characterized by incremental reward income compared to the average level To improve training efficiency and stability, a proximal policy optimization algorithm based on the Actor-Critic architecture is used to update the policy parameters.

[0057] The policy parameters are updated using a near-end policy optimization algorithm based on the Actor-Critic architecture, specifically as follows:

[0058] Old strategy Collect multiple rounds of interaction samples and calculate the advantage estimate for each time step; construct a model that includes the importance sampling ratio. objective function The difference in action selection between the new and old strategies is characterized by the importance sampling ratio, and a constraint mechanism is introduced into the objective function to limit the difference between the new and old strategies.

[0059] When importance sampling ratio When the deviation from 1 is too large, a truncation mechanism is used. right Cut to obtain ,Will Restricting the value to within between the initial agent and the next-hop agent, so as to avoid too drastic parameter update.

[0060] Finally, the parameter set is iteratively updated by using the policy gradient method, so that the new policy improves the solution performance of the incremental mixed flow transmission management problem under the premise of ensuring the stability of training:

[0061] Through the above training process, the converged multi-agent decision-making strategy is obtained

[0062]

[0063] Step 4: The trained agent is encapsulated as a multi-agent cluster and deployed in an industrial time-sensitive network to generate transmission schemes for incremental mixed flows online, specifically:

[0064] Step 4.1: When there is an incremental AVB flow or TT flow that needs to configure a transmission scheme, the transmission requirements of the flow are handed over to the initial agent for processing. The initial agent completes the first-hop transmission decision according to the flow type and the current network state, and outputs the corresponding single-hop transmission action.

[0065] Step 4.2: Determine whether the first-hop decision has formed an end-to-end transmission scheme covering the source node and the destination node. If not, transfer the decision-making right to the corresponding next-hop agent according to the next-hop link specified by the first-hop decision. The next-hop agent continues to output single-hop transmission actions under the constraints of its local observation state and historical action chain, and repeats the above process until a complete end-to-end transmission scheme is generated or it is determined that the transmission constraints of the incremental flow cannot be met under the current network conditions.

[0066] The above only describes the preferred embodiments of the present application and does not limit the idea of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.​​​

Claims

1. A method for multi-agent cooperative scheduling of incremental mixed flows in an industrial time-sensitive network, characterized in that, Comprise: Step 1: Analyze the transmission characteristics of AVB flows and TT flows in industrial time-sensitive networks, and based on network calculus theory, combine bandwidth competition and delay fluctuation in the incremental mixed flow scenario to model the dynamic interference caused by the incremental mixed flow in detail; Step 2: Combine the dynamic interference factors introduced by the incremental mixed flow, and clearly define the key transmission management variables of AVB flows and TT flows, including transmission route determination and time slot division; Step 3: Design a traffic scheduling model based on multi-agent collaboration, and based on multi-agent collaborative decision-making, generate an end-to-end transmission scheme that meets the delay and bandwidth constraints when the incremental mixed flow is accessed; Step 4: Encapsulate the trained agent as a multi-agent cluster and deploy it to the industrial time-sensitive network to generate transmission schemes for incremental mixed flows online.

2. The multi-agent collaborative scheduling method for incremental mixed flows in an industrial time-sensitive network according to claim 1, characterized in that, The topology of the industrial time-sensitive network, TSN, is represented as , is a set of network nodes, is a set of directed links: wherein, , and denote the start node, end node and bandwidth of link respectively; is the set of ports in the network, denotes the source port of link corresponding to port , and for any port , the links with this port as source are denoted by The on-demand deployment of industrial applications will introduce mixed streams containing both AVB streams and TT streams to TSN networks, which will cause An index sequence representing the incremental mixed stream is denoted as The incremental mixed stream is denoted as , which is defined as follows: wherein, represents a start point of the incremental mixed flow, represents an end point of the incremental mixed flow, represents a data size of the incremental mixed flow, represents a transmission period of the incremental mixed flow, represents a type identification of the incremental mixed flow, represents an upper limit of tolerable end-to-end delay of the incremental mixed flow; The incremental mixed flow will exacerbate the bandwidth competition among all network AVB flows or compress the available bandwidth of the AVB flow, and raise the single-hop delay of the existing AVB flow; By modeling the delay in the single-hop transmission process of the AVB flow, the dynamic interference caused by the incremental mixed flow can be effectively described; The upper limit of the single-hop delay of the AVB flow transmitted by the port is the maximum difference between the service curve and the arrival curve of the AVB flow transmitted by the port in the horizontal direction; Port Service curve for upper AVB stream transmission For: wherein, denotes the bandwidth of the link ; denotes the port time offset relative to the network start time at which the TT flow dedicated slots and the guard band are cumulatively active For the AVB stream with index in the network , suppose it enters the network through the first-hop port after being generated by the terminal, and its arrival curve at is: wherein, denotes the average generation rate over its period, denotes its possible maximum instantaneous sending amount; for any port via which it is forwarded in subsequent hops , which is the arrival curve of the output curve of the previous hop port : in, For flow In the port and its corresponding links The upper limit of single-hop delay during transmission; this value is added to the time synchronization error between nodes. Afterwards, as bias embedding exist arrival curve To depict the port Flow Output to port The effect.

3. The multi-agent collaborative scheduling method for incremental mixed flows in an industrial time-sensitive network according to claim 2, characterized in that, Said step 2 is specifically: Step 2.1: The transmission scheme for AVB streams only needs to specify the route, while the TT stream has to determine the joint route and time-slot scheduling scheme hop-by-hop, denoted as the mixed stream delta transmission scheme is defined as: in, For incremental mixed flow The number of transmission hops, for In other words, Indicates incremental mixed flow In the The links traversed during the jump, For link set The Kleene positive closure represents the set of all potential routes with a hop count of at least one; when When it is a TT stream, considering that it needs to be transmitted within one supercycle... Then, it must be allocated on each link it passes through. One time slot; Indicates in Up to be assigned to TT stream of The index sequence of time slots, family It contains all The possible sequences of time slot indices, and express and The Kleene positive closure of the Cartesian product covers all combinations of "link-slot index" union terms that contain at least one hop; each term in this closure corresponds to a potential routing-scheduling union TT stream end-to-end transmission scheme. Step 2.2: To ensure the link order connection in traffic routing and avoid forming a loop, the routing in the transmission scheme needs to meet the following constraints: wherein, denotes the start node of a link , denotes the end node of a link , denotes the end node of a link , denotes the start node of a link ; When For TT flows, the time slots allocated to it on each link further satisfy the following constraints: wherein, is the thindex in the slot scheduled for the thtransmission within a super cycle; H denotes the super cycle configuration of the TSN network, denotes the duration of a slot; the offset between each pair of adjacent slots must strictly match the stream period to guarantee the periodicity of the TT stream transmission.

4. The multi-agent collaborative scheduling method for incremental mixed flows in an industrial time-sensitive network according to claim 2, characterized in that, Said step 3 is specifically: Step 3.1: Construct a multi-agent Markov decision process; The solving of the incremental mixed flow transmission management problem is converted into a multi-agent Markov decision process, in which a state space is represented, when a transmission decision is made for a certain incremental mixed flow, an initial state quantity containing flow attribute information and network overall state is represented; an action space shared by all agents is represented; in the incremental mixed flow hop-by-hop cooperation of the agents, in the process of constructing an end-to-end transmission scheme, the output action of each agent involves the routing selection of the current hop, thereby specifying which adjacent agent the next hop is connected to for decision making, if the agents successively participate in and cooperatively complete the hop-by-hop decision process, the process is represented as: in, Indicates the decision-making process. To characterize the overall agent action of the end-to-end transmission scheme, corresponding to incremental hybrid streams end-to-end transmission scheme , For the first A smart agent The output single-hop transmission action, for the agent In other words, Defined as from the previous The action chain formed by the sequential decision-making outputs of each agent guides and constrains its current single-hop transmission decision; the overall decision of multiple agents can be represented as an action sequence composed of the single-hop actions of each agent, which corresponds to the end-to-end transmission scheme. Step 3.2: Design the state input of the multi-agent; For the An intelligent agent participating in decision-making Its state input is determined by the initial state. With action chain Composition; initial state Contains Stream Attribute information And the overall network information obtained based on topological connectivity. ; Targeting intelligent agents Action chain It is concretized into the local observation state in the local network segment it maps to. This local observation characterizes the cumulative effect of preceding decisions on adjacent links; In local observation, if the increment mixed flow is an AVB flow, the agent focuses on the single-hop upper bound of the AVB flow on the adjacent link, and evaluates the potential impact on the link load via the transmission over the link; if the increment mixed flow is a TT flow, the agent refines the observation granularity to the available time slots on the adjacent link, and analyzes the fluctuations in the link transmission load and time slot utilization introduced by the time slot allocation selection; initial state with local state complete input formed in conjunction as state input of the agent in hop-by-hop cooperation; Step 3.3: For different types of incremental flows, define differentiated action output methods for agents; For AVB flows, the agent outputs an action in state For this hop routing, only the next hop transmission link needs to be specified; while for TT flows, the agent outputs an action in state which contains both the next hop transmission link and the set of time slots allocated on this link to ensure that the routing scheme and time slot allocation scheme of TT flows are mutually consistent and conflict-free under the joint decision framework of routing and time slot scheduling. Step 3.4: To enable multi-agents to jointly optimize the end-to-end transmission performance of incremental mixed flows during collaboration, construct a reward function based on task completion; When the multi-agent collaboration successfully generates an end-to-end transmission scheme that meets the bandwidth and delay constraints for a certain incremental flow, the transmission scheme is considered as an effective configuration, and all agents participating in the decision-making process are given a positive reward, and the reward value is set to +1; If any hop in the decision-making chain cannot find a feasible next-hop link or available time slot, resulting in the inability to form a complete end-to-end transmission scheme, the collaboration is determined to fail, and all participating agents are given a negative reward, and the reward value is set to -1; Step 3.5: Configure a flow-aware neural network FANet for each agent to realize the policy mapping from state input to action output; Step 3.6: Train a multi-agent policy based on the proximal policy optimization algorithm; integrate the policy parameters of all agents into... Based on this, the overall decision-making strategy of the intelligent agent cluster is defined as follows: In the multi-agent Markov decision-making process, the agents obtain cumulative rewards by generating transmission schemes for incremental hybrid streams. And by leveraging the advantage function To measure in the initial state Take action below Characterized by incremental reward income compared to the average level To improve training efficiency and stability, a proximal policy optimization algorithm based on the Actor-Critic architecture is used to update the policy parameters.

5. The multi-agent collaborative scheduling method for incremental mixed flows in industrial time-sensitive networks according to claim 4, characterized in that, Said step 3.5 is specifically: For the initial state of the state input The flow characteristics in the state And network characteristics Two parts, through the shared flow feature extraction layer and the network feature extraction layer to process, to extract high-dimensional features related to flow attributes and overall network state; Introduce a heterogeneous feature extraction layer to adaptively process the local state corresponding to different traffic types So that the feature extraction of different dimensions Is carried out respectively under the condition that the incremental flow is AVB flow and TT flow; In​ After feature extraction is complete, the hidden features output from each feature extraction layer are concatenated, and the concatenated features are input into the corresponding decision module according to the traffic type, outputting the final action distribution or action selection; the agent... The decision-making strategy is represented as ,in for The policy parameters of a neural network.

6. The multi-agent collaborative scheduling method for incremental mixed flows in an industrial time-sensitive network according to claim 4, characterized in that, In step 3.6, the proximal policy optimization algorithm based on the Actor-Critic architecture is used to update the policy parameters, specifically: Old policy Collect multiple rounds of interaction samples, calculate the advantage estimate corresponding to each time step; construct a target function containing the importance sampling ratio of the old policy , depict the difference between the old and new policies in action selection through the importance sampling ratio, and introduce a constraint mechanism in the target function to limit the difference between the old and new policies; When importance sampling ratio When deviation 1 is too large, a truncation mechanism is adopted To Cutting gets , the value of is limited to and , so as to avoid too violent parameter update; Finally, the parameter set is updated iteratively by using the policy gradient method to make the new policy improve the solution performance of the incremental mixed streaming management problem under the premise of ensuring the stability of training. Through the above training process, the converged multi-agent decision strategy is obtained .

7. The method of claim 4, wherein, Said step 4 is specifically: Step 4.1: When there is an incremental AVB flow or TT flow that needs to be configured with a transmission scheme, the transmission requirements of the flow are handed over to the initial agent for processing, and the initial agent completes the first-hop transmission decision according to the flow type and the current network state, and outputs the corresponding single-hop transmission action; Step 4.2: Determine whether the first-hop decision has formed an end-to-end transmission scheme covering the source node and the destination node, if not completed, according to the next-hop link specified by the first-hop decision, the decision right is transferred to the corresponding next-hop agent; the next-hop agent continues to output single-hop transmission action under the constraint of its local observation state and historical action chain, and repeats the above process until a complete end-to-end transmission scheme is generated or it is determined that the transmission constraints of the incremental flow cannot be met under the current network conditions.

Citation Information

Patent Citations

  • Routing and scheduling method and device for time-triggered traffic in time-sensitive network, and readable storage medium

    CN115883438A

  • Mixed flow dynamic route planning method and device and computer equipment

    CN117135111A

  • Mixed flow-oriented time-sensitive online scheduling method in industrial Internet of Things

    CN117675680A

  • Efficient routing method for ultra-high bandwidth service flow in time-sensitive network

    CN119030914A

  • Industrial internet-oriented network slice resource allocation method and related equipment

    CN119814685A

Cited By

  • Method and device for cross-domain mixed flow routing and scheduling of deterministic network

    CN122179358A

  • A method and apparatus for deterministic network inter-domain mixed flow routing and scheduling

    CN122179358B