Multi-agent scheduling method, device and equipment of industrial network and storage medium

By dividing the industrial network into multiple sub-synchronization domains and constructing constraints and observable Markov decision processes, the deterministic transmission problem of multi-form business flows in complex industrial scenarios is solved, improving the cross-domain collaborative scheduling efficiency and transmission accuracy of chained flows.

CN116566981BActive Publication Date: 2026-03-24BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

How to build deterministic network technologies with "timely and accurate" performance characteristics to effectively support the converged and shared transmission of multi-form service flows and adapt to the deterministic transmission needs of multi-form service flows in complex industrial scenarios.

Method used

The single synchronization domain in the industrial network is divided into multiple sub-synchronization domains, and a network model is constructed using multiple sub-synchronization domains. By deploying a single domain controller, a set of constraints and a partially observable Markov decision process are constructed, and a cross-domain collaborative scheduling strategy for chained business flows is formulated.

Benefits of technology

It enables deterministic transmission of multi-form business flows in complex industrial scenarios, reduces resource occupancy and scheduling failure rate, and improves the deterministic transmission quality of the Industrial Internet.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116566981B_ABST
    Figure CN116566981B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer software, and discloses a multi-agent scheduling method, device and equipment of an industrial network and a storage medium, the method comprising the following steps: dividing a single synchronization domain in the industrial network into multiple sub-synchronization domains, and constructing a network model by using the multiple sub-synchronization domains, wherein each sub-synchronization domain is controlled by a corresponding single-domain controller; constructing a constraint condition set with the aim of reducing resource occupancy and scheduling failure rate; constructing a partially observable Markov decision process according to the constraint condition set and the network model, and the partially observable Markov decision process is used for formulating an optimal strategy for cross-domain collaborative scheduling of a chain service flow. The multi-agent scheduling method of the industrial network provided by the application can effectively adapt to a complex industrial scene and meet the determined transmission requirements of multi-form service flows.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer software technology, and more specifically to a multi-agent scheduling method, apparatus, equipment, and storage medium for industrial networks. Background Technology

[0002] The Industrial Internet is helping the manufacturing industry move towards networking and intelligence. With the deep integration of information technology for management and operational technology for production, industrial scenarios are constantly giving rise to many diverse types of business flows, such as classic business flows centered on time-triggered flows, complex chained flows characterized by "sequential triggering of sub-flows," and intelligent business flows for new industrial applications. Therefore, how to build deterministic network technologies with "timely and accurate" performance characteristics to effectively support the converged and shared transmission of multi-form business flows is a major challenge currently facing the Industrial Internet.

[0003] Therefore, there is an urgent need for a service flow scheduling method that can adapt to complex industrial scenarios and meet the deterministic transmission requirements of multi-form service flows. Summary of the Invention

[0004] In view of this, the present invention provides a multi-agent scheduling method, apparatus, device and storage medium for industrial networks, which can adapt to complex industrial scenarios and effectively meet the deterministic transmission requirements of multi-form service flows.

[0005] In a first aspect, the present invention provides a multi-agent scheduling method for industrial networks, comprising:

[0006] A single synchronization domain in an industrial network is divided into multiple sub-synchronization domains, and a network model is constructed using these multiple sub-synchronization domains, wherein each sub-synchronization domain is controlled by a corresponding single-domain controller.

[0007] To reduce resource utilization and scheduling failure rate, a set of constraints is constructed;

[0008] Based on the set of constraints and the network model, a partially observable Markov decision process is constructed, which is used to formulate the optimal strategy for cross-domain collaborative scheduling of chained business flows.

[0009] The multi-agent scheduling method for industrial networks provided by this invention controls each sub-synchronization domain by setting a corresponding single-domain controller and constructing a partially observable Markov decision process to formulate the optimal strategy for cross-domain collaborative scheduling of chained business flows. This method can effectively adapt to complex industrial scenarios and meet the deterministic transmission requirements of multi-form business flows.

[0010] In one optional implementation, dividing a single synchronization domain in the industrial network into multiple sub-synchronization domains includes:

[0011] A single synchronization domain is divided into multiple sub-synchronization domains based on the controller location, network size, and access method.

[0012] The multi-agent scheduling method for industrial networks provided by this invention can effectively ensure the flexibility and scalability of chained service flow scheduling by dividing a single synchronization domain according to the controller location, network size and access method.

[0013] In one alternative implementation, the set of constraints, aimed at reducing resource utilization and scheduling failure rate, includes:

[0014]

[0015] Among them, E m,n,t This represents the cumulative energy consumption value. The maximum threshold for cumulative energy consumption is represented by m, n, M, N, and R. The set of intelligent agent numbers and the set of access device numbers are also represented by m, n, M, and N. The time is represented by t and R.

[0016]

[0017] Among them, T m,n,t This represents the cumulative time cost value. The maximum threshold for cumulative time overhead is represented by m, n, M, N, and R.

[0018]

[0019] Among them, T t T represents the current period delay value. max This represents the maximum periodic delay, where t represents time and R represents the set of times.

[0020]

[0021] Among them, O m,n,t The node selection value is 1, which means that the corresponding node is selected to participate in the current learning cycle, and 0 means that the corresponding node is not selected to participate in the current learning cycle; m represents the agent number, n represents the access device number, M represents the set of agent numbers, N represents the set of access device numbers, t represents time, and R represents the set of time.

[0022]

[0023]

[0024]

[0025] in, For downlink spectrum resource allocation decision variables. For the decision variables of uplink spectrum resource allocation, and η m,n,t For onboard computing resource decision variables.

[0026] The multi-agent scheduling method for industrial networks provided by this invention establishes a set of constraints aimed at reducing resource occupancy and scheduling failure rate by satisfying requirements such as learning cycle delay, cumulative time budget, and long-term energy consumption of nodes. This method can effectively adapt to complex industrial scenarios and thus meet the deterministic transmission requirements of multi-form business flows.

[0027] In one optional implementation, constructing a partially observable Markov decision process based on the set of constraints and the network model includes:

[0028] One agent is deployed in each single-domain controller to acquire sub-synchronization domain state information associated with the target agent and to generate corresponding scheduling decisions based on the sub-synchronization domain state information.

[0029] The multi-agent scheduling method for industrial networks provided by this invention generates corresponding scheduling decisions by acquiring the sub-synchronization domain state information associated with the agents, thereby achieving the goal of formulating the optimal strategy for cross-domain collaborative scheduling of chained flows.

[0030] In one optional implementation, acquiring the sub-synchronization domain state information associated with the target agent includes:

[0031]

[0032] Among them, S m,t Represents status information, T m,n,t E represents the latency requirement of a chained stream. m,n,t Let m represent the remaining energy of the node, n represent the agent number, M represent the set of agent numbers, and N represent the set of access device numbers.

[0033] In one alternative implementation, the method includes:

[0034] After the agent executes the scheduling decision, the agent obtains a corresponding reward value, which is used to evaluate the effectiveness of the scheduling decision.

[0035] In one optional implementation, the step of the agent obtaining a corresponding reward value after executing the scheduling decision includes:

[0036] If the agent successfully schedules the data after executing the scheduling decision, a reward value is given to the agent; the reward value indicates that the scheduling decision satisfies the constraints.

[0037] If the agent fails to schedule after executing the scheduling decision, a penalty value is given to the agent; the penalty value indicates that the scheduling decision does not meet the constraints.

[0038] The multi-agent scheduling method for industrial networks provided by this invention evaluates the effectiveness of decisions made by agents based on the state of their associated sub-synchronization domains. This reward mechanism penalizes decisions that fail to meet the deterministic latency requirements of chained flows and rewards decisions that do meet these requirements, thereby optimizing scheduling decisions.

[0039] Secondly, the present invention provides a multi-agent scheduling device for industrial networks, comprising:

[0040] The synchronization domain partitioning module is used to divide a single synchronization domain in an industrial network into multiple sub-synchronization domains and to construct a network model using the multiple sub-synchronization domains, wherein each sub-synchronization domain is controlled by a corresponding single-domain controller.

[0041] The set building module is used to construct a set of constraints with the goal of reducing resource utilization and scheduling failure rate;

[0042] The process construction module is used to construct a partially observable Markov decision process based on the set of constraints and the network model. The partially observable Markov decision process is used to formulate the optimal strategy for cross-domain collaborative scheduling of chained business flows.

[0043] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the multi-agent scheduling method for industrial networks described in the first aspect or any corresponding embodiment thereof.

[0044] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the multi-agent scheduling method for an industrial network according to the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0045] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0046] Figure 1 This is a flowchart illustrating a multi-agent scheduling method for an industrial network according to an embodiment of the present invention.

[0047] Figure 2 This is a schematic diagram regarding a single synchronization domain according to an embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of multiple sub-synchronization domains according to an embodiment of the present invention;

[0049] Figure 4 This is a schematic diagram illustrating the generation of scheduling decisions based on state information according to an embodiment of the present invention;

[0050] Figure 5 This is a schematic diagram of a multi-agent cooperative scheduling algorithm according to an embodiment of the present invention;

[0051] Figure 6 This is a structural block diagram of a multi-agent scheduling device for an industrial network according to an embodiment of the present invention;

[0052] Figure 7 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] The Industrial Internet is helping the manufacturing industry move towards networking and intelligence. With the deep integration of information technology for management and operational technology for production, industrial scenarios are constantly giving rise to many diverse business flow types, such as classic business flows centered on time-triggered flows, complex chain flows characterized by "sequential triggering of sub-flows," and intelligent business flows for new industrial applications. Therefore, how to construct deterministic network technologies with "timely and accurate" performance characteristics to effectively support the converged and shared transmission of multi-form business flows is a major challenge facing the Industrial Internet. In recent years, artificial intelligence, due to its dynamic perception and flexible decision-making characteristics, has been initially explored in network resource management and transmission scheduling optimization. Compared with traditional scheduling methods based on optimization theory and heuristic algorithms, intelligent algorithms not only possess global and rapid decision-making capabilities but also can flexibly adapt to complex industrial scenarios with dynamically changing traffic volumes, effectively meeting the deterministic transmission needs of multi-form business flows. Therefore, it is necessary to propose a new scheduling scheme utilizing artificial intelligence technology, which can make business flow scheduling more concise, timely, and flexible.

[0055] In view of this, embodiments of the present invention provide a multi-agent collaborative scheduling method and apparatus for chained business flows. This invention addresses the specific needs of complex industrial chained flow scheduling by designing a multi-industry controller collaborative scheduling architecture. By deploying dedicated controllers in each subdomain, a multi-domain collaborative scheduling network model is formed, thereby flexibly supporting deterministic cross-domain transmission of chained flows. System modeling is performed to address the sequential triggering characteristics of chained flows, forming a multi-objective optimization problem that minimizes bandwidth occupancy and scheduling failure rate. Furthermore, the performance boundary of the scheduling strategy is studied based on node size. Finally, multi-agent collaborative learning theory is used to solve the collaborative scheduling of "sequentially triggered" business sub-flows spanning multiple synchronization domains, effectively supporting the collaborative transmission of "monitor-control-execution" chained flows and improving the deterministic transmission quality of the Industrial Internet.

[0056] According to an embodiment of the present invention, a multi-agent scheduling method for an industrial network based on chained business flows is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0057] This embodiment provides a multi-agent scheduling method for industrial networks. Figure 1 This is a flowchart of a multi-agent scheduling method for industrial networks according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:

[0058] Step S101: Divide the single synchronization domain in the industrial network into multiple sub-synchronization domains, and construct a network model using the multiple sub-synchronization domains, wherein each sub-synchronization domain is controlled by a corresponding single-domain controller.

[0059] Specifically, the centrally managed single synchronization domain is divided into multiple sub-synchronization domains, thereby constructing a multi-synchronization fusion network model. A whole-domain partitioning method is used, dividing the original entire domain into multiple time-synchronized sub-domains, i.e., multiple sub-synchronization domains. Here, the original entire domain refers to the single synchronization domain, and synchronization refers to time synchronization, with the controller centrally controlling the operation of all devices according to a time cycle. Each sub-synchronization domain is controlled by a corresponding single-domain controller, which can be a distributed controller.

[0060] More specifically, a centrally managed single synchronization domain, such as Figure 2 As shown: a single global controller controls all production units; the resulting multiple sub-synchronization domains are as follows: Figure 3 As shown: Each sub-synchronization domain is controlled by a corresponding single-domain controller, and each sub-synchronization domain may contain one or more production units.

[0061] In some preferred embodiments, step S101 further includes step S1011, which is as follows: The single synchronization domain is divided according to the controller location, network size, and access method to obtain multiple sub-synchronization domains. By dividing the single synchronization domain according to the controller location, network size, and access method to obtain multiple sub-synchronization domains, the flexibility and scalability of chained flow scheduling can be effectively guaranteed.

[0062] Specifically, the controller location, such as Figure 3 The location of domain controllers 1 and 2; network size, which can be understood as the number of production units contained in the divided sub-synchronization domains, whether it is one production unit or multiple production units, for example. Figure 3 In this context, a production unit is considered as a sub-synchronization domain; access methods can be categorized into wired network access and wireless network access.

[0063] Step S102: To reduce resource utilization and scheduling failure rate, construct a set of constraints.

[0064] Specifically, in order to meet the requirements of learning cycle latency, cumulative time budget, and long-term energy consumption of nodes, a resource scheduling optimization problem is established with the goal of reducing resource occupancy and scheduling failure rate. The objective function is as follows:

[0065]

[0066] Where, r UR Indicates resource utilization rate, d URξ represents the scheduling failure rate, 'o' represents the node selection value, and ξ represents the node selection value. ↓ Let ξ represent the decision variable for downlink spectrum resource allocation. ↑ η represents the decision variable for uplink spectrum resource allocation, and η represents the decision variable for onboard computing resources.

[0067] The set of constraints for the multi-objective resource scheduling optimization problem, namely:

[0068]

[0069] Among them, E m,n,t This represents the cumulative energy consumption value. The maximum threshold for cumulative energy consumption is represented by m, n, M, N, and R. The set of intelligent agent numbers and the set of access device numbers are also represented by m, n, M, and N. The time is represented by t and R.

[0070]

[0071] Among them, T m,n,t This represents the cumulative time cost value. The maximum threshold for cumulative time overhead is represented by m, n, M, N, and R.

[0072]

[0073] Among them, T t T represents the current period delay value. max This represents the maximum periodic delay, where t represents time and R represents the set of times.

[0074]

[0075] Among them, O m,n,t The node selection value is 1, which means that the corresponding node is selected to participate in the current learning cycle, and 0 means that the corresponding node is not selected to participate in the current learning cycle; m represents the agent number, n represents the access device number, M represents the set of agent numbers, N represents the set of access device numbers, t represents time, and R represents the set of time.

[0076]

[0077]

[0078]

[0079] in, For downlink spectrum resource allocation decision variables. For the decision variables of uplink spectrum resource allocation, and η m,n,t For onboard computing resource decision variables.

[0080] Step S103: Based on the set of constraints and the network model, construct a partially observable Markov decision process, which is used to formulate the optimal strategy for cross-domain collaborative scheduling of chained business flows.

[0081] Specifically, the multi-objective resource scheduling optimization problem can be transformed into a partially observable Markov decision process involving multiple agents. In a network model constructed from multiple sub-synchronous domains, the multiple agents are deployed in various single-domain controllers.

[0082] In some preferred embodiments, step S103 includes steps S1031-S1032, as follows:

[0083] Step S1031: Deploy an agent in each single-domain controller to obtain sub-synchronization domain state information associated with the target agent, and generate corresponding scheduling decisions based on the sub-synchronization domain state information.

[0084] Specifically, each agent can and can only acquire the service flow and network state information of the associated single synchronization domain, which can be represented by the following formula:

[0085]

[0086] Among them, S m,t Represents status information, T m,n,t E represents the latency requirement of a chained stream. m,n,t Let m represent the remaining energy of the node, n represent the agent number, M represent the set of agent numbers, and N represent the set of access device numbers.

[0087] Specifically, after obtaining the state information of the sub-synchronization domains associated with the agent, a parameterized strategy can be used to characterize the dynamic characteristics of the system, thereby formulating an optimal strategy for cross-domain collaborative scheduling of chained business flows; this parameterized strategy can be based on a set of chained flow state information. Mapping the corresponding cross-domain scheduling decision The specific formula is shown below:

[0088]

[0089] Among them, a m,t Represents cross-domain scheduling decisions, η m,n,t This represents the decision variable for onboard computing resources, and its value range is [0,1].

[0090] More specifically, such as Figure 4 As shown: The decision network Actorπ1 formulates a scheduling decision a1 based on the state of the single synchronization domain corresponding to the agent; after the agent executes the corresponding scheduling decision a1, the state of the single synchronization domain is transferred to the new state.

[0091] Step S1032: After the agent executes the scheduling decision, the agent obtains the corresponding reward value, which is used to evaluate the effectiveness of the scheduling decision.

[0092] Specifically, when the agent executes the mapped cross-domain scheduling decision Then, a reward function can be used to... To measure the effectiveness of the decision.

[0093] In some optional embodiments, step S1032 further includes steps a1 and a2:

[0094] Step a1: If the agent successfully schedules the data after executing the scheduling decision, a reward value is given to the agent; the reward value indicates that the scheduling decision satisfies the constraints.

[0095] Step a2: If the agent fails to schedule after executing the scheduling decision, a penalty value is given to the agent; the penalty value indicates that the scheduling decision does not meet the constraints.

[0096] Specifically, when agent m dictates the state t of the associated sub-synchronization domain... m,s Make scheduling decision a m,t Subsequently, the agent receives a corresponding reward value, which is used to evaluate the effectiveness of the decisions made. To minimize the evaluation loss of hierarchical federated learning, the reward function of the multi-agent resource adaptation algorithm is defined as:

[0097]

[0098] Since multiple agents collaborate to make decisions on cross-domain scheduling of the chained stream, +1 is used as the reward value for each agent. -1 represents a penalty factor used to punish decisions that fail to meet the constraints, i.e., decisions made by the algorithm fail to meet the deterministic latency requirements of the chained stream.

[0099] In some optional embodiments, the method further includes:

[0100] Step S104: Based on the multi-agent cooperative scheduling algorithm, multiple agents are embedded into the multi-synchronization domain network model to perform distributed network management and control.

[0101] The above step S104 specifically includes the following steps S1041-S1043, such as... Figure 5 As shown:

[0102] Step S1041: Initialize algorithm parameters to initialize the weight parameters of the policy network and evaluation network of each agent and instantiate the experience storage pool at the same time.

[0103] Specifically, after the algorithm training begins, the weight parameters θ of the policy network and evaluation network of each agent can be initialized. m φ m,1 φ m,2 , Simultaneously instantiate the experience storage pool K1;

[0104] Where, θ m These are the weight parameters of the Actor neural network (i.e., the policy network); φ m,1 and φ m,2 These are the weight parameters of the target networks Critic1 and Critic2; These are the weight parameters for evaluating the networks Critic1 and Critic2.

[0105] Step S1042: Establish experience samples to provide usable training samples for algorithm parameter updates and decision performance optimization.

[0106] Specifically, a multi-dimensional tuple is constructed to represent the agent's experience samples, including state, decision, reward, and next state, as follows:

[0107]

[0108] Among them, S m,t Indicates the current status information, a m,t Represents the scheduling decision, r m,t S' represents the reward received. m,t Indicates the next state.

[0109] The above multidimensional tuples can be constructed through the following steps:

[0110] First, the agents perceive the state information S of the service flow from their respective associated synchronization domains. m,t ;

[0111] Secondly, based on the perceived state information, the agent uses a policy network... Formulate scheduling decisions a m,t ; where a m,t ={η m,t} represents the scheduling decision for cross-domain collaborative transmission of intelligent agent m.

[0112] Then, the agents execute the corresponding decisions and receive the corresponding rewards r. m,t At this point, the states of each synchronization domain will transition to the next state S'.m,t During the state transition, after the agent makes a decision, the remaining energy of the node will change, thus completing the state transition.

[0113] Finally, the tuples formed by sampling Store in experience storage pool K r This provides usable training samples for algorithm parameter updates and decision performance optimization.

[0114] Step S1043: Update the algorithm parameters to complete the collection of empirical samples and the updating of neural network parameters.

[0115] Specifically, during the algorithm update phase, multiple agents collaborate to solve a partially observable Markov decision process, alternately collecting experience samples and updating the parameters of the fully connected neural network. For agent m, during each update, it needs to retrieve data from the experience storage pool K. r Extract K from b A sample of experiences To update the parameters.

[0116] Step S1043 above also includes the following steps b1-b4, as follows:

[0117] Step b1: Update the evaluation network.

[0118] Specifically, the edge agent can minimize the mean square error function L(φ) m,v The parameters of the evaluation network are updated independently. It can be expressed by the following formula:

[0119]

[0120]

[0121] Wherein, L(φ) m,v ) represents the target Q value, and K represents the set of current states and transition states of all agents, respectively; b This indicates the number of samples drawn from the experience sample pool; Indicate decision-making, ) represents the target value, and k represents the sample number. Indicates a reward;

[0122] log(.) is a function that calculates the entropy value of resource adaptation decisions. and φ represents the "state-decision" Q-values ​​calculated by the current evaluation network and the target evaluation network, respectively. To effectively mitigate the bias problem in the policy improvement process, two evaluation networks, φ and φ', are used in the current and target networks, respectively. m,v , Each evaluation network can be computed with a Q-value. When calculating the mean squared error in the second part, only the minimum of the two Q-values ​​is used. Furthermore, the evaluation network parameters are updated using gradient descent, i.e.:

[0123]

[0124] in, It is the gradient of the mean squared error function; the other parameters have been explained in the above formula.

[0125] Step b2, update the policy network.

[0126] Specifically, the policy gradient method is used to update the parameters of the policy network, and its objective function is defined as:

[0127]

[0128] Where J is the objective function of the Actor network. Let α represent the decision made by M agents in the t-th decision cycle. m Indicates the weighting factor;

[0129] Among them, Gaussian noise is added to the input of the policy network. t It enables the reparameterization of policy networks. This design is advantageous for obtaining a lower variance estimate. Therefore, the objective function can be rewritten as:

[0130]

[0131] Furthermore, the update gradient of the policy network parameters can be calculated, i.e.:

[0132]

[0133] Step b3: Update the decision entropy weight coefficients.

[0134] Specifically, the weighting coefficient between the decision entropy term and the agent's reward function will be α. m The parameterization is performed as a fully connected neural network, and parameter updates are used to dynamically adjust the weight coefficients in the maximum entropy objective function. α can be achieved using the following formula. m Iterative optimization, i.e.

[0135]

[0136] in, H represents the set of decisions of all edge agents, H′ represents the target entropy value, and the other parameters have been explained in the above formula.

[0137] Furthermore, according to the updated target network, the proposed algorithm employs a soft parameter update method, using the current evaluation network's weight parameters φ in a certain proportion to make the learning process smoother. m,v Update the weight parameters φ′ of the target evaluation network m,v ,Right now

[0138]

[0139] Where τ∈(0,1) represents the update scaling factor of the target evaluation network relative to the current evaluation network, and φ m,v φ′ represents the weight parameters of the current evaluation network. m,v This indicates that the weight parameters of the target evaluation network are being updated.

[0140] The technical solution provided by this invention has the following effects:

[0141] This invention provides a multi-agent cooperative scheduling method and apparatus for chained service flows. By dividing the system into multiple synchronization domains, a multi-objective optimization problem and a set of constraints are constructed and transformed into a partially observable Markov decision process. A parameterized strategy is used to characterize the dynamic characteristics of the system, thus formulating an optimal strategy for cross-domain cooperative scheduling of chained flows. Furthermore, a multi-agent cooperative scheduling algorithm is designed to solve the cooperative scheduling of service sub-flows that are "triggered sequentially" across multiple synchronization domains, effectively supporting the cooperative transmission of "monitor-control-execution" chained flows and improving the deterministic transmission quality of the Industrial Internet.

[0142] This embodiment also provides a multi-agent scheduling device for an industrial network, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0143] This embodiment provides a multi-agent scheduling device for industrial networks, such as... Figure 6 As shown, it includes:

[0144] The synchronization domain partitioning module 501 is used to divide a single synchronization domain in an industrial network into multiple sub-synchronization domains and to construct a network model using the multiple sub-synchronization domains, wherein each sub-synchronization domain is controlled by a corresponding single-domain controller.

[0145] Set construction module 502 is used to construct a set of constraints with the goal of reducing resource utilization and scheduling failure rate;

[0146] The process construction module 503 is used to construct a partially observable Markov decision process based on the set of constraints and the network model. The partially observable Markov decision process is used to formulate the optimal strategy for cross-domain collaborative scheduling of chained business flows.

[0147] In some optional implementations, the synchronization domain partitioning module 501 includes:

[0148] The partitioning unit is used to divide a single synchronization domain into multiple sub-synchronization domains according to the controller location, network size, and access method.

[0149] In some alternative implementations, the set of constraints, aimed at reducing resource utilization and scheduling failure rate, includes:

[0150]

[0151] Among them, E m,n,t This represents the cumulative energy consumption value. The maximum threshold for cumulative energy consumption is represented by m, n, M, N, and R. The set of intelligent agent numbers and the set of access device numbers are also represented by m, n, M, and N. The time is represented by t and R.

[0152]

[0153] Among them, T m,n,t This represents the cumulative time cost value. The maximum threshold for cumulative time overhead is represented by m, n, M, N, and R.

[0154]

[0155] Among them, T t T represents the current period delay value. max This represents the maximum periodic delay, where t represents time and R represents the set of times.

[0156]

[0157] Among them, O m,n,tThe node selection value is 1, which means that the corresponding node is selected to participate in the current learning cycle, and 0 means that the corresponding node is not selected to participate in the current learning cycle; m represents the agent number, n represents the access device number, M represents the set of agent numbers, N represents the set of access device numbers, t represents time, and R represents the set of time.

[0158]

[0159]

[0160]

[0161] in, For downlink spectrum resource allocation decision variables. For the decision variables of uplink spectrum resource allocation, and η m,n,t For onboard computing resource decision variables.

[0162] In some alternative implementations, process construction module 503 includes:

[0163] The agent deployment unit is used to deploy one agent in each single-domain controller, to obtain sub-synchronization domain state information associated with the target agent, and to generate corresponding scheduling decisions based on the sub-synchronization domain state information.

[0164] In some optional implementations, the agent deployment unit is further configured to obtain sub-synchronization domain state information associated with the target agent using the following formula:

[0165]

[0166] Among them, S m,t Represents status information, T m,n,t E represents the latency requirement of a chained stream. m,n,t Let m represent the remaining energy of the node, n represent the agent number, M represent the set of agent numbers, and N represent the set of access device numbers.

[0167] In some alternative embodiments, the apparatus further includes:

[0168] The reward module is used to provide a corresponding reward value to the agent after the agent executes the scheduling decision. The reward value is used to evaluate the effectiveness of the scheduling decision.

[0169] In some alternative implementations, the reward module includes:

[0170] A reward unit is configured to grant a reward value to the agent if the agent successfully executes the scheduling decision; the reward value indicates that the scheduling decision satisfies the constraints.

[0171] A penalty unit is used to give the agent a penalty value if the agent fails to schedule after executing the scheduling decision; the penalty value indicates that the scheduling decision does not meet the constraints.

[0172] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0173] In this embodiment, the multi-agent scheduling device for the industrial network is presented in the form of functional units. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0174] This invention also provides a computer device having the above-described features. Figure 7 The multi-agent scheduling device for the industrial network shown.

[0175] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 7 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 7 Take a processor 10 as an example.

[0176] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0177] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0178] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0179] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0180] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0181] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0182] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A multi-agent scheduling method for industrial networks, characterized in that, include: A single synchronization domain in an industrial network is divided into multiple sub-synchronization domains, and a network model is constructed using these multiple sub-synchronization domains, wherein each sub-synchronization domain is controlled by a corresponding single-domain controller. The method of dividing a single synchronization domain in an industrial network into multiple sub-synchronization domains includes: A single synchronization domain is divided into multiple sub-synchronization domains based on the controller location, network size, and access method. To reduce resource utilization and scheduling failure rate, a set of constraints is constructed; Based on the set of constraints and the network model, a partially observable Markov decision process is constructed, which is used to formulate the optimal strategy for cross-domain collaborative scheduling of chained business flows. The set of constraints, aimed at reducing resource utilization and scheduling failure rate, includes: in, This represents the cumulative energy consumption value. This indicates the maximum threshold for cumulative energy consumption. Indicates the agent's serial number. Indicates the serial number of the access device. A set representing the indices of intelligent agents. This represents the set of access device serial numbers. Indicates time, A set representing time; in, This represents the cumulative time cost value. This represents the maximum threshold for cumulative time cost. Indicates the agent's serial number. Indicates the serial number of the access device. A set representing the indices of intelligent agents. This represents the set of access device serial numbers. Indicates time, A set representing time; in, This indicates the current period delay value. This indicates the maximum periodic delay. Indicates time, A set representing time; in, This represents the node selection value. A node selection value of 1 indicates that the corresponding node is selected to participate in the current learning cycle, while a node selection value of 0 indicates that the corresponding node is not selected to participate in the current learning cycle. Indicates the agent's serial number. Indicates the serial number of the access device. A set representing the indices of intelligent agents. This represents the set of access device serial numbers. Indicates time, A set representing time; in, For downlink spectrum resource allocation decision variables. Decision variables for uplink spectrum resource allocation, and For onboard computing resource decision variables.

2. The method according to claim 1, characterized in that, The construction of a partially observable Markov decision process based on the set of constraints and the network model includes: One agent is deployed in each single-domain controller to acquire sub-synchronization domain state information associated with the target agent and to generate corresponding scheduling decisions based on the sub-synchronization domain state information.

3. The method according to claim 2, characterized in that, The acquisition of sub-synchronization domain state information associated with the target agent includes: in, Indicates status information, This indicates the latency requirements of a chained stream. Indicates the remaining energy of the node. Indicates the agent's serial number. Indicates the serial number of the access device. A set representing the indices of intelligent agents. This represents the set of access device serial numbers. Indicates time, A set representing time.

4. The method according to claim 3, characterized in that, The method includes: After the agent executes the scheduling decision, the agent obtains a corresponding reward value, which is used to evaluate the effectiveness of the scheduling decision.

5. The method according to claim 4, wherein after the agent executes the scheduling decision, the agent obtains the corresponding reward value, comprising: If the agent successfully schedules the task after executing the scheduling decision, the agent is given a reward value. The reward value indicates that the scheduling decision meets the constraints. If the agent fails to schedule after executing the scheduling decision, a penalty value is given to the agent; the penalty value indicates that the scheduling decision does not meet the constraints.

6. A multi-agent scheduling device for an industrial network, characterized in that, The device includes: The synchronization domain partitioning module is used to divide a single synchronization domain in an industrial network into multiple sub-synchronization domains and to construct a network model using the multiple sub-synchronization domains, wherein each sub-synchronization domain is controlled by a corresponding single-domain controller. The synchronization domain partitioning module is specifically used to partition a single synchronization domain according to the controller location, network size, and access method, resulting in multiple sub-synchronization domains. The set building module is used to construct a set of constraints with the goal of reducing resource utilization and scheduling failure rate; The process construction module is used to construct a partially observable Markov decision process based on the set of constraints and the network model. The partially observable Markov decision process is used to formulate the optimal strategy for cross-domain collaborative scheduling of chained business flows. The set construction module is specifically used for in, This represents the cumulative energy consumption value. This indicates the maximum threshold for cumulative energy consumption. Indicates the agent's serial number. Indicates the serial number of the access device. A set representing the indices of intelligent agents. This represents the set of access device serial numbers. Indicates time, A set representing time; in, This represents the cumulative time cost value. This represents the maximum threshold for cumulative time cost. Indicates the agent's serial number. Indicates the serial number of the access device. A set representing the indices of intelligent agents. This represents the set of access device serial numbers. Indicates time, A set representing time; in, This indicates the current period delay value. This indicates the maximum periodic delay. Indicates time, A set representing time; in, This represents the node selection value. A node selection value of 1 indicates that the corresponding node is selected to participate in the current learning cycle, while a node selection value of 0 indicates that the corresponding node is not selected to participate in the current learning cycle. Indicates the agent's serial number. Indicates the serial number of the access device. A set representing the indices of intelligent agents. This represents the set of access device serial numbers. Indicates time, A set representing time; in, For downlink spectrum resource allocation decision variables. Decision variables for uplink spectrum resource allocation, and For onboard computing resource decision variables.

7. A computer device, characterized in that, include: A memory and a processor are interconnected, the memory stores computer instructions, and the processor executes the computer instructions to perform the multi-agent scheduling method for an industrial network as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the multi-agent scheduling method for an industrial network as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Collaborative optimization scheduling method, device and equipment for multiple virtual power plants, and storage medium

    CN114036825A

  • DDQN-based TSN routing method

    CN116132353A