Trusted scheduling method and apparatus for dynamic blockchain-based industrial wireless network
By constructing a dynamic blockchain system and a multi-agent learning algorithm, the task and resource scheduling of industrial wireless networks is optimized, solving the problem of low trust between edge servers and industrial equipment, and realizing an efficient and secure computing environment.
Patent Information
- Application Number
- PCT/CN2025/079447
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-02
- Filing Date
- 2025-02-27
- Publication Date
- 2026-02-05
AI Technical Summary
In industrial wireless networks, the low trust between edge servers and industrial devices makes edge computing prone to failure, and existing technologies cannot effectively solve this problem.
An industrial wireless network system based on dynamic blockchain is constructed. It adopts a multi-agent Markov decision process model and a round-robin multi-agent deep reinforcement learning algorithm to dynamically adjust the leader edge server and the blockchain timeout window, optimize the joint scheduling of tasks and resources, and ensure security and trustworthiness.
It improves the reliable processing efficiency of tasks in industrial wireless network systems, enhances the system's flexibility and adaptability, and ensures the security and reliability of computing.
Smart Images

Figure CN2025079447_05022026_PF_FP_ABST
Abstract
Description
A dynamic blockchain-based industrial wireless network trusted scheduling method and device TECHNICAL FIELD
[0001] The present application relates to the transmission scheduling technology of the industrial internet, in particular to a dynamic blockchain-based industrial wireless network trusted scheduling method and device. BACKGROUND
[0002] With the rapid development of the industrial internet, more and more industrial devices are connected to the internet, and people, machines and things are connected to each other, realizing a highly intelligent industrial system. However, due to the limited computing frequency of industrial devices and the limited battery capacity, running computationally intensive applications on these devices can cause significant computational delay and excessive energy consumption, seriously reducing user experience. To solve this problem, edge computing has emerged, allowing tasks to be offloaded to edge servers for real-time processing.
[0003] However, due to the access of a large number of heterogeneous industrial devices and the open wireless network environment, the trustworthiness between industrial devices and edge servers in the industrial internet may be affected, and data and other information are easily stolen by attackers. Therefore, blockchain technology is introduced to enhance its security and privacy.
[0004] However, simply adding blockchain to industrial wireless networks may not guarantee the trustworthiness between edge servers and industrial devices, and may even exacerbate the overhead, leading to the failure of edge computing. SUMMARY
[0005] Therefore, the present application provides a dynamic blockchain-based industrial wireless network trusted scheduling method and device, mainly to solve the problem of low trustworthiness between edge servers and industrial devices in the current industrial wireless network task and resource joint scheduling process, and the problem of easy failure of edge computing.
[0006] To solve the above problems, the present application provides a dynamic blockchain-based industrial wireless network trusted scheduling method, comprising:
[0007] Constructing an industrial wireless network system based on a dynamic blockchain mechanism;
[0008] Based on the preset constraint conditions corresponding to the task and resource joint scheduling stage of the industrial wireless network system and the consensus stage of the dynamic blockchain mechanism, a model is constructed to maximize the task trusted processing efficiency, and an optimization model for scheduling the industrial wireless network system is obtained, which carries each to be optimized parameter in the optimization model;
[0009] Based on a preset multi-agent Markov decision process model, the optimization model is restructured to obtain a target optimization model;
[0010] The target optimization model is optimized based on observation information of each industrial device in the industrial wireless network system collected in real time, a preset round-robin multi-agent deep reinforcement learning algorithm model is used to optimize the target optimization model, and target parameter values corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system are obtained.
[0011] To solve the above problems, the application provides an industrial wireless network trusted scheduling device based on a dynamic blockchain, comprising:
[0012] An industrial wireless network system construction module is configured to construct an industrial wireless network system based on a dynamic blockchain mechanism.
[0013] A model construction module is configured to construct a model based on preset constraint conditions corresponding to a task and resource joint scheduling phase of the industrial wireless network system and a consensus phase of the dynamic blockchain mechanism, and to maximize the task trusted processing efficiency, so as to obtain an optimization model for scheduling the industrial wireless network system, wherein each to-be-optimized parameter is carried in the optimization model.
[0014] A model reconstruction module is configured to reconstruct the optimization model based on a preset multi-agent Markov decision process model, so as to obtain a target optimization model.
[0015] An optimization module is configured to optimize the target optimization model based on observation information of each industrial device in the industrial wireless network system collected in real time, and to obtain target parameter values corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system.
[0016] To solve the above problems, the application provides a storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the above-mentioned industrial wireless network trusted scheduling method based on a dynamic blockchain.
[0017] To solve the above problems, the application provides an electronic device, which at least includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the above-mentioned industrial wireless network trusted scheduling method based on a dynamic blockchain when executing the computer program stored in the memory.
[0018] The beneficial effects in the present application: the present application constructs an industrial wireless network system based on a dynamic blockchain mechanism; based on the preset constraint conditions corresponding to the task and resource joint scheduling stage of the industrial wireless network system and the consensus stage of the dynamic blockchain mechanism, the model is constructed with the goal of maximizing the task trusted processing efficiency, and an optimization model for scheduling the industrial wireless network system is obtained, which carries each to be optimized parameter; the dynamic blockchain mechanism is adopted, the leader edge server and the blockchain timeout window are dynamically adjusted according to the network state and the task demand, which not only guarantees the security and credibility, but also improves the flexibility and adaptability of the system. Based on the preset multi-agent Markov decision process model, the optimization model is restructured to obtain a target optimization model; based on the observation information of each industrial device in the industrial wireless network system collected in real time, a preset round-robin multi-agent deep reinforcement learning algorithm model is used to optimize the target optimization model, and the target parameter value corresponding to each to be optimized parameter for scheduling the industrial wireless network system is obtained, which improves the trusted processing efficiency of the task in the industrial wireless network system.
[0019] The above description is only a summary of the technical scheme of the present application. In order to enable the technical means of the present application to be more clearly understood, the following specific embodiments of the present application are described in detail in accordance with the contents of the specification, and in order to enable the above and other purposes, characteristics and advantages of the present application to be more obvious and easy to understand, the following specific embodiments of the present application are described in detail. BRIEF DESCRIPTION OF DRAWINGS
[0020] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become apparent to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be limiting on the present application. Moreover, the same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0021] Fig. 1 shows a flowchart of an industrial wireless network trusted scheduling method based on a dynamic blockchain according to an embodiment of the present application;
[0022] Fig. 2 shows a flowchart of an industrial wireless network trusted scheduling method based on a dynamic blockchain according to another embodiment of the present application;
[0023] Fig. 3 shows a structural diagram of an industrial wireless network system according to an embodiment of the present application;
[0024] Fig. 4 shows a structure diagram of a preset round-robin multi-agent deep reinforcement learning algorithm model according to an embodiment of the present application;
[0025] Fig. 5 shows a structural block diagram of an industrial wireless network trusted scheduling device based on a dynamic blockchain according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] The various methods as well as the features of the application herein will be described with reference to the drawings.
[0027] It is to be understood that various modifications can be made to the embodiments described herein. Thus, the description is not to be considered as limiting, but merely as a description of exemplary embodiments. Other modifications of the application will occur to those skilled in the art upon reading the description of the application.
[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the application and, together with the general description of the application given above, and the detailed description of the embodiments given below, serve to explain the principles of the present application.
[0029] These and other characteristics of the present application will become apparent from the following description of the preferred forms given, by way of non-limiting example only, with reference to the attached drawings.
[0030] It is also to be understood that even though a few specific embodiments of the present application have been described, many other modifications will occur to those skilled in the art.
[0031] The above and other aspects, features, and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which:
[0032] Specific embodiments of the present application are described hereinafter, by way of non-limiting example only; however, it should be understood that many other configurations of the present application will fall within the scope of the present application. Accordingly, the application is not limited to the specific embodiments described herein, but includes all possible embodiments thereof that can be defined with the scope of the appended claims.
[0033] The specification can use phrases such as "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", which can refer to one or more of the same or different embodiments of the application.
[0034] The embodiments of the present application provide a dynamic blockchain-based industrial wireless network trusted scheduling method, as shown in FIG. 1, which includes:
[0035] Step S101: Construct an industrial wireless network system based on a dynamic blockchain mechanism;
[0036] In the implementation process, the industrial wireless network system includes a plurality of edge servers and a plurality of industrial devices, the edge servers are powered by a power grid and are connected by wires, the plurality of edge servers include an edge computing server, a base station, and a blockchain node server; the edge computing server provides edge computing and content caching services; the base station provides communication services; the blockchain node server provides security for communication and computing services; the industrial devices generate tasks and perform local computing on the tasks, and support migration of the tasks to the edge servers through a wireless channel for edge computing; the industrial devices, such as a robot vision sensor with a computing-intensive task, are connected to the edge servers through a wireless connection and can offload tasks to the edge servers for collaborative computing; the dynamic blockchain mechanism, each edge server acts as a blockchain node and is responsible for recording transaction information of task offloading and computing. Here, the edge servers play two roles: leaders and followers. The leaders are responsible for verifying transactions and generating blocks, while the followers are responsible for initiating transactions and storing blocks. And the blockchain timeout window and the leader edge server will change dynamically with the environment.
[0037] Step S102: Based on the preset constraint conditions corresponding to the task and resource joint scheduling phase of the industrial wireless network system and the consensus phase of the dynamic blockchain mechanism, respectively, a model is constructed to maximize the task trusted processing efficiency, and an optimization model for scheduling the industrial wireless network system is obtained, the optimization model carries each to be optimized parameter;
[0038] In the implementation process, each of the preset constraint conditions includes a task offloading ratio constraint condition, a bandwidth allocation ratio constraint condition, a local computing frequency constraint condition, an energy consumption constraint condition, a blockchain security constraint condition, a computing frequency allocation constraint condition, a task deadline constraint condition, and a trustworthiness constraint condition. Each of the to-be-optimized parameters includes a task division ratio, a communication bandwidth allocation ratio, a computing frequency of an edge server for offloading task processing and block generation, a dynamic consensus waiting window in a blockchain, and a leader agent elected by an edge server.
[0039] Step S103: Based on a preset multi-agent Markov decision process model, the optimization model is reconstructed to obtain a target optimization model;
[0040] In the specific implementation process of this step, the agent set, the agent observation state set, and the agent action set of the preset multi-agent Markov decision process model are constructed; a function is constructed based on a preset trusted computing reward function, a preset timeout penalty function, and a preset consensus penalty function to obtain a reward function of the preset multi-agent Markov decision process model; the model is constructed based on the agent set, the agent observation state set, the agent action set, and the reward function to obtain the preset multi-agent Markov decision process model; and the optimization model of the industrial wireless network scheduling is model-converted based on the preset multi-agent Markov decision process model and the reward function to obtain a target optimization model.
[0041] Step S104: The target optimization model is optimized based on the real-time collected observation information of each industrial device in the industrial wireless network system using a preset round-robin multi-agent deep reinforcement learning algorithm model to obtain target parameter values corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system.
[0042] In the specific implementation process of this step, the observation information of each industrial device in the industrial wireless network system obtained in real time is used to perform state prediction using a preset round-robin multi-agent deep reinforcement learning algorithm to obtain target parameter values corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system; each target parameter value is used to perform calculation based on the target optimization model to obtain a target reward value; the target parameter values are executed to obtain target observation information of each industrial device in the next state; each observation information, each data target parameter value, the target reward value, and the target observation information are input into a preset experience database; and a foundation is laid for subsequent offline updating of the preset round-robin multi-agent deep reinforcement learning algorithm model in specific applications.
[0043] The application constructs an industrial wireless network system based on a dynamic blockchain mechanism; preset constraint conditions corresponding to a task and resource joint scheduling phase of the industrial wireless network system and a consensus phase of the dynamic blockchain mechanism are used to construct a model with the goal of maximizing the task trusted processing efficiency, an optimization model for scheduling the industrial wireless network system is obtained, and each to-be-optimized parameter is carried in the optimization model; the dynamic blockchain mechanism is used to dynamically adjust the leader edge server and the blockchain timeout window according to the network state and the task demand, which not only guarantees the security and trustworthiness, but also improves the flexibility and adaptability of the system. The optimization model is restructured based on a preset multi-agent Markov decision process model to obtain a target optimization model; the target optimization model is optimized based on a preset round-robin multi-agent deep reinforcement learning algorithm model using the observation information of each industrial device in the industrial wireless network system collected in real time to obtain target parameter values corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system, thereby improving the trusted processing efficiency of the task in the industrial wireless network system.
[0044] Another embodiment of the application provides another dynamic blockchain-based industrial wireless network trusted scheduling method, as shown in FIG. 2, which includes the following steps:
[0045] Step S201: Construct an industrial wireless network system based on a dynamic blockchain mechanism;
[0046] In the specific implementation process, as shown in FIG. 3, it is a structural schematic diagram of the industrial wireless network system of the application, which includes M edge servers and N industrial devices, and sets M={1, 2, …, M} and N={1, 2, …, N} respectively. The edge servers are powered by the power grid and are connected by wires, and the plurality of edge servers include edge computing servers, base stations, and blockchain node servers; the edge computing servers provide edge computing and content caching services; the base stations provide communication services; the blockchain node servers provide security assurance for communication and computing services; the industrial devices generate tasks and perform local computing on the tasks, and support migrating tasks to edge servers for edge computing through wireless channels; the industrial devices, such as robot vision sensors with computing-intensive tasks, are connected to edge servers through wireless connections and can offload tasks to edge servers for collaborative computing; the dynamic blockchain mechanism, each edge server acts as a blockchain node, responsible for recording transaction information of task offloading and computing. Here, the edge servers play two roles: leader and follower. The leader is responsible for verifying transactions and generating blocks, while the follower is responsible for initiating transactions and storing blocks. And the blockchain timeout window and the leader edge server will change dynamically with the environment.
[0047] Step S202: Construct the preset constraints;
[0048] In the specific implementation process of this step, constraints are constructed based on the proportion of tasks offloaded from each industrial device to each edge server, resulting in the task offload ratio constraint C1 for the task and resource joint scheduling stage; the mathematical expression of the task offload ratio constraint can be expressed as the following formula (1):
[0049] Where, x n,m x represents the proportion of tasks offloaded from industrial equipment n to edge server m. n,0 This indicates the proportion of tasks computed locally. When x... n,m When x = 1, it means that the nth industrial device will offload all tasks to the mth edge server for full edge computing, while when x = 1... n,m When = 0, it means that the nth industrial device does not offload any tasks to the mth edge server.
[0050] Based on the bandwidth ratio allocated to each of the aforementioned edge servers by each of the aforementioned industrial devices, constraints are constructed to obtain the bandwidth allocation ratio constraint C2 for the task and resource joint scheduling phase; the mathematical expression of the bandwidth allocation ratio constraint can be expressed as the following formula (2):
[0051] Where, r n,m This represents the bandwidth ratio allocated by the m-th edge server to the tasks generated by the n-th industrial device.
[0052] Constraints are constructed based on the predetermined maximum computing frequency parameters corresponding to each of the industrial devices, resulting in local computing frequency constraints C3 corresponding to each of the industrial devices; the mathematical expression of the local computing frequency constraints can be expressed as shown in the following formula (3):
[0053] Among them, f n This represents the local computing frequency used by the nth industrial device, and should not exceed the maximum computing frequency of the nth industrial device.
[0054] Constraints are constructed based on the computational energy consumption, task transmission energy consumption, and maximum battery capacity of each of the aforementioned industrial devices, resulting in energy consumption constraint C4 corresponding to each of the aforementioned industrial devices; the mathematical expression of the energy consumption constraint can be expressed as shown in the following formula (4):
[0055] The main energy consumption of the nth industrial device consists of computing energy consumption and task transmission energy consumption, which should not exceed the maximum battery capacity of the industrial device.
[0056] The constraint condition is constructed based on the predetermined success task count function of the dynamic blockchain mechanism, the preset Byzantine fault tolerance coefficient, the identification parameter of each edge server completing the offloading task of each industrial equipment, and the task proportion of each industrial equipment offloaded to each edge server, to obtain a blockchain security constraint condition C5 of the consensus stage. The mathematical expression of the blockchain security constraint condition of the consensus stage can be shown in the following formula (5):
[0057] The blockchain security constraint should meet the Byzantine fault tolerance mechanism; wherein Count() is a predetermined success task count function, used to calculate the number of z n,m =1; z n,m ∈{-1,1} represents whether the mth edge server completes the task of offloading the nth industrial equipment within the task deadline, and p is the Byzantine fault tolerance coefficient, that is, the constraint indicates that the number of transactions received in time to complete the task within the allowed maximum consensus delay should not be lower than the first preset threshold, is a rounding up symbol, indicating rounding a decimal value to the nearest integer value not less than the original value.
[0058] The constraint condition is constructed based on the maximum calculation frequency of each edge server, the binary indication variable of whether each edge server is a leader, and the calculation frequency allocated to block generation when each edge server acts as a leader, to obtain a calculation frequency allocation constraint condition C6 corresponding to each edge server. The mathematical expression of the calculation frequency allocation constraint condition can be shown in the following formula (6):
[0059] wherein, is the maximum calculation frequency of the mth edge server, l m ={0,1} is a binary leader indication variable, assuming that the leader edge server is m * , and represents the calculation frequency allocated for task calculation; represents the calculation frequency allocated to block generation when acting as a leader;
[0060] The constraint condition is constructed based on the predetermined task deadline of each industrial equipment, to obtain a task deadline constraint condition C7 corresponding to each industrial equipment. The mathematical expression of the task deadline constraint condition can be shown in the following formula (7):
[0061] wherein, a deadline of a task of the nth industrial equipment, i.e. a maximum tolerable processing delay of the industrial equipment.
[0062] A constraint condition is constructed based on the trust score corresponding to each edge server and a predetermined threshold value, to obtain a trustworthiness constraint condition C8 corresponding to each edge server. The mathematical expression of the trustworthiness constraint condition can be shown in the following formula (8): C8: v m >V th , m e M (8)
[0063] wherein, for the task completed by the mth edge server, the trust score v m calculated based on the number of successfully completed tasks should be greater than a set threshold value V th .
[0064] Step S203: constructing a task trust processing efficiency function;
[0065] In the implementation process, the function is constructed based on a preset trustworthiness verification delay function, a preset task offloading transmission delay function and a preset edge computing delay function, to obtain a trusted computing process function; specifically, the function is constructed based on the data amount of the verified block, the communication rate function between each industrial equipment and each edge server, to obtain the trustworthiness verification delay function The mathematical expression of the trustworthiness verification delay function can be shown in the following formula (9):
[0066] wherein, VB is the data amount of the verified block; R n,m is a communication rate function, and the mathematical expression of the communication rate function can be shown in the following formula (10):
[0067] wherein, r n,m represents the bandwidth proportion allocated by the edge server m to the industrial equipment n, B m represents the bandwidth of the edge server m, p n is the transmission power of the nth industrial equipment, h n,m is the channel power gain between the nth industrial equipment and the mth edge server, and N0 is the noise power spectral density.
[0068] The function is constructed based on the task proportion of each industrial equipment offloaded to each edge server, the task data amount of each industrial equipment offloaded to each edge server and the communication rate function between each industrial equipment and each edge server, to obtain the preset task offloading transmission delay function The mathematical expression of the preset task offloading transmission delay function can be shown in the following formula (11):
[0069] Among them, D n The amount of task data that is offloaded from the industrial equipment n to the edge server m.
[0070] Based on the proportion of tasks offloaded from each industrial device to each edge server, the computation frequency required for each task of the industrial device, and the computation frequency allocated by the edge server to the industrial device, a function is constructed to obtain the preset edge computing latency function. The mathematical expression of the preset edge computing delay function can be expressed as shown in the following formula (12):
[0071] Among them, C n This represents the computational frequency required for the task of industrial equipment n.
[0072] The trusted computing process function The mathematical expression for can be represented by the following formula (13):
[0073] Based on the trusted computing process function, the preset transaction record report delay function, the preset maximum allowed consensus waiting time function, and the preset maximum actual consensus waiting time function, a function is constructed to obtain the actual consensus waiting time function; specifically, it is based on the transaction record data volume TR and the transmission rate between the edge server. Perform function construction to obtain the preset transaction record report delay function. The mathematical expression for the preset transaction record reporting delay function can be expressed as follows (14):
[0074] Based on the trusted computing process function The preset transaction record report delay function and dynamic consensus waiting window Perform function construction to obtain the maximum allowed consensus waiting time function. The mathematical expression for the maximum allowed consensus waiting time function can be expressed as follows (15):
[0075] Based on the trusted computing process function and the preset transaction record report delay function The function is constructed to obtain the actual maximum consensus waiting time function. The mathematical expression for the maximum consensus waiting time function can be expressed as follows (16):
[0076] based on the allowed maximum consensus waiting time function and the actual maximum consensus waiting time function function construction is performed to obtain the actual consensus waiting time function The mathematical expression of the actual consensus waiting time function can be shown in the following formula (17):
[0077] based on the actual consensus waiting time function and a preset block generation delay function function construction is performed to obtain an edge trusted computing delay function The mathematical expression of the edge trusted computing delay function can be shown in the following formula (18):
[0078] The mathematical expression of the preset block generation delay function can be shown in the following formula (19):
[0079] wherein, is the computing frequency allocated by the leader for block generation.
[0080] based on the edge trusted computing delay function and a local delay function corresponding to each of the industrial devices function construction is performed to obtain a task trusted computing delay function T corresponding to each of the industrial devices n . The mathematical expression of the task trusted computing delay function can be shown in the following formula (20):
[0081] based on the task trusted computing delay function T corresponding to the same industrial device n and the task data volume D n function construction is performed to obtain the trusted processing efficiency function U corresponding to the same industrial device n . The mathematical expression of the trusted processing efficiency function can be shown in the following formula (21):
[0082] Step S204: Based on the preset constraint conditions corresponding to the task and resource joint scheduling phase of the industrial wireless network system and the consensus phase of the dynamic blockchain mechanism respectively, a model is constructed with the goal of maximizing the task trusted processing efficiency to obtain an optimization model for scheduling the industrial wireless network system;
[0083] In the specific implementation process of this step, the mathematical expression of the optimization model can be shown in the following formula (22):
[0084] wherein the optimization model represents maximizing the task trustful processing efficiency, X, R, F, T, L are the set of parameters to be optimized, X = {x n,m} M×N is the set of task partitioning ratios, R = {r n,m} M×N is the set of communication bandwidth allocation ratios, is the set of computing frequencies allocated by the edge server for task computation and block generation, is the set of dynamic consensus waiting windows in the blockchain, L = {l m} M is the set of leaders elected by all edge servers.
[0085] Step S205: constructing a preset multi-agent Markov decision process model;
[0086] In the implementation process, the agent set, the agent observation state set, and the agent action set of the preset multi-agent Markov decision process model are constructed. Specifically, the agent set is M = {1, 2, …, M}, that is, each edge server is regarded as an agent, learns its own optimal action strategy by observing the environment state, and cooperates with other agents to achieve the maximum task trustful processing efficiency. The agent observation state set includes the task state, the industrial equipment state, and the edge server state. At each decision time t, the task state includes the task size D n (t), the required computing frequency C n (t), and the task deadline T The industrial equipment state includes the transmission power p n (t), the channel power gain h n,m (t), the battery capacity B , and the device local computing frequency f The edge server state includes the edge computing frequency f , the bandwidth B m (t), and the trust score v m (t). Therefore, the entire environment state can be represented by the following formula (23):
[0087] wherein by observing the environment state, each agent will obtain its own state, denoted as S m (t). The action set is the set of action strategies taken by the agent, and the mathematical expression of the action strategy set can be represented by the following formula (24):
[0088] wherein x relates to the task partitioning ratio.n,m (t), communication bandwidth allocation r n,m (t), computing frequency allocation and consensus latency window and leader election l m (t).
[0089] The mathematical expression of the preset trusted computing reward function can be shown in the following formula (25):
[0090] wherein, is the penalty of computing timeout, is the penalty of consensus failure, and α and β are weight factors of the penalties. The penalty of computing timeout The mathematical expression of the penalty of computing timeout
[0091] The penalty of consensus failure The mathematical expression of the penalty of consensus failure
[0092] Based on the preset trusted computing reward function, the preset timeout penalty function and the preset consensus penalty function, a function is constructed to obtain a reward function of the preset multi-agent Markov decision process model. The mathematical expression of the reward function can be shown in the following formula (28):
[0093] Based on the agent set, the agent observation state set, the agent action set and the reward function, a model is constructed to obtain the preset multi-agent Markov decision process model.
[0094] Step S206: Based on the preset multi-agent Markov decision process model, the optimization model is restructured to obtain a target optimization model.
[0095] In the specific implementation process, the mathematical expression of the target optimization model can be shown in the following formula (29):
[0096] That is, under the condition of meeting the constraints, the long-term cumulative reward is maximized, that is, the trusted processing efficiency of the task is maximized.
[0097] Step S207: A preset round-robin multi-agent deep reinforcement learning algorithm model is constructed.
[0098] The step is preset in the specific implementation process, as shown in Figure 4, the rotating multi-agent deep reinforcement learning algorithm model structure diagram of the application, the initial model corresponding to each edge server is constructed, the initial model includes an initial actor neural network, two initial critic neural networks and two initial target critic neural networks;
[0099] The initial model corresponding to the predetermined initial leader agent in each edge server is initialized to obtain the initial model parameters corresponding to the initial leader agent;
[0100] Based on each historical experience data, each initial model parameter, a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function, the model training is performed on each initial model to obtain a target deep reinforcement learning model corresponding to each edge server, so that the preset rotating multi-agent deep reinforcement learning algorithm model is obtained; specifically, in the first decision period, based on the initial calculation frequency, the initial trust score and the initial channel state of each edge server, the leader election is performed to obtain an initial leader agent; the initial trust score v m The calculation mathematical formula of the initial trust score v
[0101] Wherein, z n,m ∈{1,-1} represents whether the mth edge server completes the unloading task of the nth industrial equipment on time, 1 for completion, -1 for not completed.
[0102] Based on the preset first loss function, the preset second loss function, the preset third loss function, and the preset entropy regularization loss function, the initial model corresponding to the initial leader agent is trained to obtain a current first model corresponding to the initial leader agent; specifically, based on a group of randomly extracted historical experience data, the first loss function is used for calculation and processing to obtain the first model parameter of the first initial critic neural network; based on another group of randomly extracted historical experience data, the second loss function is used for calculation and processing to obtain the second model parameter of the second initial critic neural network; the first model parameter is obtained by minimizing the calculation and processing of the first loss function; the second model parameter is obtained by minimizing the calculation and processing of the second loss function; the calculation mathematical formula of the first loss function L Q (ω1) can be shown as follows formula (31):
[0103] Wherein, ω1 is the first model parameter; Ω(t) = {S(t), A(t), R(t), S(t+1)} is λ sets of the historical experience data, wherein S(t) is the observation state at t moment, that is, the state of the whole system; A(t) is the action policy at t moment; R(t) is the reward value obtained by executing the action at t moment; S(t+1) is the observation state of the whole industrial wireless network system at the next moment after executing the action at t moment. is the first initial critic neural network; μ is an entropy regularization coefficient; is the expected value calculation function; γ is a discount factor. The second loss function L Q The calculation mathematical formula of (ω2) can be shown in the following formula (32):
[0104] Wherein: ω2 is the second model parameter; is the second initial critic neural network.
[0105] The first model parameter and the second model parameter are screened to obtain a target model parameter; the first model parameter and the second model parameter are compared to determine the smaller model parameter as the target model parameter. Based on the target model parameter, the historical experience data are calculated and processed by a third loss function of an initial actor neural network to obtain a third model parameter θ when the loss value of the third loss function is the smallest. The mathematical expression of the third loss function can be shown in the following formula (33):
[0106] Wherein, θ is the third model parameter; π θ is an initial actor neural network; f θ is a reparameterization function for action sampling. ε is a noise random variable satisfying a policy Gaussian distribution G.
[0107] Based on the historical experience data, a preset entropy regularization loss function is calculated and processed to obtain a regularization coefficient when the loss value of the preset entropy regularization loss function is the smallest. The mathematical expression of the preset entropy regularization loss function can be shown in the following formula (34):
[0108] Wherein, H0 is a target entropy;
[0109] updating the first initial critic neural network based on the first model parameter to obtain a first current critic neural network; updating the second initial critic neural network based on the second model parameter to obtain a second current critic neural network; updating the initial actor neural network based on the third model parameter to obtain a current actor model; and updating the preset entropy regularization loss function based on the regularization coefficient to obtain a current entropy regularization loss function. Specifically, each model can be updated by using a soft update function, and a mathematical expression of the soft update function can be shown in the following formula (35):
[0110] wherein τ ∈ [0, 1] is a parameter update rate.
[0111] The first current critic neural network, the second current critic neural network, the current actor model, and the current entropy regularization loss function are iteratively updated until the reward function value obtained by optimizing the target optimization model of industrial wireless network scheduling using the state parameter values obtained by state prediction of each historical experience data using the updated current actor model meets a preset condition, and the current first model is obtained. The model parameters of the current first model are distributed to each first edge server other than the initial leader agent, so as to update the initial model corresponding to each first edge server based on the model parameters, and obtain the current first model corresponding to each first edge server. The model parameters include the first model parameter, the second model parameter, the third model parameter, and the first entropy regularization coefficient corresponding to the initial leader agent. In a non-first decision period, a current leader agent is obtained by re-selecting a leader based on the current computing frequency, the current trust score, and the current channel state of each edge server. The current first model corresponding to the current leader agent is trained based on a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function, so as to update the current first model. The model parameters of the updated current first model are distributed to each second edge server other than the current leader agent, so as to update the current first model corresponding to each second edge server based on the model parameters, and the preset round-robin multi-agent deep reinforcement learning algorithm model is obtained by iterative circulation until the algorithm converges. The target deep reinforcement learning model includes a target actor neural network, each critic neural network, and a target critic neural network corresponding to each critic neural network for stabilizing the critic neural network.
[0112] Step S208: based on the observation information of each industrial device in the industrial wireless network system collected in real time, a preset round-robin multi-agent deep reinforcement learning algorithm model is used to optimize the target optimization model to obtain target parameter values corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system.
[0113] In the specific implementation process, the observation information of each industrial device in the industrial wireless network system collected in real time is subjected to state prediction by using a preset round-robin multi-agent deep reinforcement learning algorithm model, target parameter values meeting each observation information are obtained; an optimization model for scheduling the industrial wireless network system is calculated based on each target parameter value, a target reward value is obtained; target observation information of each industrial device in the next state is obtained by executing the target parameter values; each observation information, each data target parameter value, the target reward value and the target observation information are entered into a preset experience database; and a foundation is laid for subsequent offline update of the preset round-robin multi-agent deep reinforcement learning algorithm model in specific application.
[0114] The end-edge collaborative trusted computing process adopted by the present application, as shown in FIG. 4, mainly includes the following steps:
[0115] Step one, the edge server performs dynamic leader agent election and establishes a trusted blockchain;
[0116] Step two, the industrial device verifies the trustworthiness of the target edge server for task offloading, and performs task offloading for end-edge collaborative trusted computing;
[0117] Step three, the follower agent reports the transaction record to the leader agent and waits for consensus;
[0118] Step four, the leader agent verifies the block and broadcasts the verified block to update the blockchain.
[0119] The present application performs dynamic leader election and dynamic consensus waiting mechanism in a trusted edge computing environment. The leader edge server and the blockchain timeout window can be dynamically adjusted according to the network state and task demand, which not only guarantees safety and trustworthiness, but also improves the flexibility and adaptability of the system. Considering the task size and deadline, communication, computation and energy resources, and trustworthiness constraints, the task allocation, communication and computation frequency allocation, leader election and consensus waiting window are jointly optimized, so that the scheduling of tasks and resources can be dynamically adjusted according to the changes of the scene, ensuring safety and trustworthiness while efficiently completing task computation. The preset round-robin training distributed execution architecture is adopted, which improves the trusted processing efficiency of tasks in the industrial wireless network system.
[0120] Another embodiment of the present application provides an industrial wireless network trusted scheduling device based on dynamic blockchain, as shown in FIG. 5, which includes:
[0121] An industrial wireless network system construction module 1 for constructing an industrial wireless network system based on a dynamic blockchain mechanism;
[0122] A model construction module 2 is configured to construct a model for maximizing task trusted processing efficiency based on preset constraint conditions corresponding to a task and resource joint scheduling phase of the industrial wireless network system and a consensus phase of the dynamic blockchain mechanism respectively, to obtain an optimization model for scheduling the industrial wireless network system, and the optimization model carries each to be optimized parameter;
[0123] A model reconstruction module 3 is configured to reconstruct the optimization model based on a preset multi-agent Markov decision process model to obtain a target optimization model.
[0124] An optimization module 4 is configured to optimize the target optimization model based on real-time collected observation information of each industrial device in the industrial wireless network system by using a preset round-robin multi-agent deep reinforcement learning algorithm model, to obtain target parameter values corresponding to each to be optimized parameter for scheduling the industrial wireless network system.
[0125] In the specific implementation process, the dynamic blockchain-based industrial wireless network trusted scheduling device further comprises: each preset constraint condition construction module, which is specifically configured to: construct a task offloading ratio constraint condition of the task and resource joint scheduling phase based on a task proportion of each industrial device offloaded to each edge server; construct a bandwidth allocation ratio constraint condition of the task and resource joint scheduling phase based on a bandwidth proportion of each edge server allocated to each industrial device; construct a local computing frequency constraint condition corresponding to each industrial device based on a predetermined maximum computing frequency parameter corresponding to each industrial device; construct an energy consumption constraint condition corresponding to each industrial device based on a computing energy consumption, a task transmission energy consumption and a maximum battery capacity corresponding to each industrial device; construct a blockchain security constraint condition of the consensus phase based on a predetermined successful task processing count function of the dynamic blockchain mechanism, a preset Byzantine fault tolerance coefficient, an identifier parameter of each edge server completing offloaded tasks of each industrial device, and a task proportion of each industrial device offloaded to each edge server; construct a computing frequency allocation constraint condition corresponding to each edge server based on a maximum computing frequency of each edge server, a binary indicator variable of whether each edge server is a leader, and a computing frequency allocated to block generation when each edge server acts as a leader; construct a task deadline constraint condition corresponding to each industrial device based on a predetermined task deadline of each industrial device; and construct a trustworthiness constraint condition corresponding to each edge server based on a trust score of a task successful completion amount corresponding to each edge server and a predetermined threshold.
[0126] In the implementation process, the industrial wireless network trusted scheduling device based on a dynamic blockchain further includes a trusted processing efficiency function construction module, which is specifically configured to: perform function construction based on a preset trust degree verification time delay function, a preset task offloading transmission time delay function, and a preset edge computing time delay function to obtain a trusted computing process function; perform function construction based on the trusted computing process function, a preset transaction record reporting time delay function, a preset allowed maximum consensus waiting time function, and a preset actual maximum consensus waiting time function to obtain an actual consensus waiting time delay function; perform function construction based on the actual consensus waiting time delay function and a preset block generation time delay function to obtain an edge trusted computing time delay function; perform function construction based on the edge trusted computing time delay function and a local time delay function corresponding to each industrial device to obtain a task trusted computing time delay function corresponding to each industrial device; and perform function construction based on the task trusted computing time delay function corresponding to a same industrial device and a task data volume to obtain the trusted processing efficiency function corresponding to the same industrial device.
[0127] In the implementation process, the industrial wireless network trusted scheduling device based on a dynamic blockchain further includes a preset multi-agent Markov decision process model construction module, which is specifically configured to: construct an agent set, an agent observation state set, and an agent action set of the preset multi-agent Markov decision process model; perform function construction based on a preset trusted computing reward function, a preset timeout penalty function, and a preset consensus penalty function to obtain a reward function of the preset multi-agent Markov decision process model; and perform model construction based on the agent set, the agent observation state set, the agent action set, and the reward function to obtain the preset multi-agent Markov decision process model.
[0128] In the specific implementation process, the industrial wireless network trusted scheduling device based on a dynamic blockchain further includes a preset round-robin multi-agent deep reinforcement learning algorithm model construction module, which is specifically configured to: construct an initial model corresponding to each edge server, the initial model including an initial actor neural network, two initial critic neural networks, and two initial target critic neural networks; initialize the initial model corresponding to a predetermined initial leader agent in each edge server to obtain an initial model parameter corresponding to the initial leader agent; perform model training on each initial model based on each historical experience data, each initial model parameter, a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function to obtain a target deep reinforcement learning model corresponding to each edge server, so as to obtain the preset round-robin multi-agent deep reinforcement learning algorithm model; wherein the target deep reinforcement learning model includes a target actor neural network, each critic neural network, and a target critic neural network corresponding to each critic neural network for stabilizing the critic neural network.
[0129] In the implementation process, the preset round-robin multi-agent deep reinforcement learning algorithm model construction module is further configured to: in a first decision period, perform leader election based on initial computing frequencies, initial trust scores, and initial channel states of the edge servers to obtain an initial leader agent; perform model training on an initial model corresponding to the initial leader agent based on a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function to obtain a current first model corresponding to the initial leader agent; distribute model parameters of the current first model to each first edge server that is not the initial leader agent, so as to perform model updating on initial models corresponding to the first edge servers based on the model parameters, and obtain current first models corresponding to the first edge servers, wherein the model parameters include first model parameters, second model parameters, third model parameters, and a first entropy regularization coefficient corresponding to the initial leader agent; in a non-first decision period, perform leader election again based on current computing frequencies, current trust scores, and current channel states of the edge servers to obtain a current leader agent; perform model training on a current first model corresponding to the current leader agent based on the preset first loss function, the preset second loss function, the preset third loss function, and the preset entropy regularization loss function, so as to update the current first model; distribute model parameters of the updated current first model to each second edge server that is not the current leader agent, so as to perform model updating on current first models corresponding to the second edge servers based on the model parameters, and perform cyclic iteration until the algorithm converges, so as to train the preset round-robin multi-agent deep reinforcement learning algorithm model.
[0130] In the implementation process, the optimization module 4 is specifically configured to: based on observation information of each industrial device in the industrial wireless network system collected in real time, perform state prediction on a target actor neural network in a preset round-robin multi-agent deep reinforcement learning algorithm model to obtain target parameter values corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system; the to-be-optimized parameters include: a task division ratio, a communication bandwidth allocation ratio, a computing frequency of an edge server for offloading task processing and block generation, a dynamic consensus waiting window in a block chain, and a leader agent of edge server election.
[0131] The application constructs an industrial wireless network system based on a dynamic blockchain mechanism; preset constraint conditions corresponding to a task and resource joint scheduling phase of the industrial wireless network system and a consensus phase of the dynamic blockchain mechanism are used to construct a model with the goal of maximizing task trusted processing efficiency, an optimization model for scheduling the industrial wireless network system is obtained, and each to-be-optimized parameter is carried in the optimization model; a dynamic blockchain mechanism is used to dynamically adjust the leader edge server and the blockchain timeout window according to the network state and the task demand, which not only guarantees safety and credibility, but also improves the flexibility and adaptability of the system. The optimization model is reconstructed based on a preset multi-agent Markov decision process model to obtain a target optimization model; a preset round-robin multi-agent deep reinforcement learning algorithm model is used to optimize the target optimization model based on real-time collected observation information of each industrial device in the industrial wireless network system, target parameter values corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system are obtained, and the trusted processing efficiency of the task in the industrial wireless network system is improved.
[0132] Another embodiment of the application provides a storage medium storing a computer program, which is executed by a processor to implement the following method steps:
[0133] Step one, constructing an industrial wireless network system based on a dynamic blockchain mechanism;
[0134] Step two, constructing a model with the goal of maximizing task trusted processing efficiency based on preset constraint conditions corresponding to a task and resource joint scheduling phase of the industrial wireless network system and a consensus phase of the dynamic blockchain mechanism, obtaining an optimization model for scheduling the industrial wireless network system, and carrying each to-be-optimized parameter in the optimization model;
[0135] Step three, reconstructing the optimization model based on a preset multi-agent Markov decision process model to obtain a target optimization model;
[0136] Step four, optimizing the target optimization model based on a preset round-robin multi-agent deep reinforcement learning algorithm model based on real-time collected observation information of each industrial device in the industrial wireless network system, obtaining target parameter values corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system.
[0137] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit, module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units or modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0138] The specific implementation process of the above method steps can refer to the embodiments of any of the above industrial wireless network trusted scheduling methods based on a dynamic blockchain mechanism, which will not be repeated here.
[0139] The present application constructs an industrial wireless network system based on a dynamic blockchain mechanism; based on the preset constraint conditions corresponding to the task and resource joint scheduling stage of the industrial wireless network system and the consensus stage of the dynamic blockchain mechanism, an optimization model for scheduling the industrial wireless network system is obtained by model construction with the goal of maximizing the task trusted processing efficiency, and each to-be-optimized parameter is carried in the optimization model; the dynamic blockchain mechanism is adopted, and the leader edge server and the blockchain timeout window are dynamically adjusted according to the network state and the task demand, which not only guarantees the safety and trustworthiness, but also improves the flexibility and adaptability of the system. Based on the preset multi-agent Markov decision process model, the optimization model is restructured to obtain a target optimization model; based on the real-time collected observation information of each industrial device in the industrial wireless network system, a preset round-robin multi-agent deep reinforcement learning algorithm model is used to optimize the target optimization model, and the target parameter value corresponding to each to-be-optimized parameter for scheduling the industrial wireless network system is obtained, thereby improving the trusted processing efficiency of the task in the industrial wireless network system.
[0140] Another embodiment of the present application provides an electronic device, which at least includes a memory and a processor, the memory stores a computer program, and the processor implements the following method steps when executing the computer program on the memory:
[0141] Step one, constructing an industrial wireless network system based on a dynamic blockchain mechanism;
[0142] Step two, based on the preset constraint conditions corresponding to the task and resource joint scheduling stage of the industrial wireless network system and the consensus stage of the dynamic blockchain mechanism, an optimization model for scheduling the industrial wireless network system is obtained by model construction with the goal of maximizing the task trusted processing efficiency, and each to-be-optimized parameter is carried in the optimization model;
[0143] Step three, based on the preset multi-agent Markov decision process model, the optimization model is restructured to obtain a target optimization model;
[0144] Step four, based on the observation information of each industrial device in the industrial wireless network system collected in real time, a preset round-robin multi-agent deep reinforcement learning algorithm model is used to optimize the target optimization model, and target parameter values corresponding to each optimization parameter for scheduling the industrial wireless network system are obtained.
[0145] The specific implementation process of the above method steps can refer to the embodiments of any of the above industrial wireless network trusted scheduling methods based on a dynamic blockchain. The embodiments will not be repeated here.
[0146] The present application constructs an industrial wireless network system based on a dynamic blockchain mechanism. Based on the preset constraint conditions corresponding to the task and resource joint scheduling stage of the industrial wireless network system and the consensus stage of the dynamic blockchain mechanism, a model is constructed to maximize the task trusted processing efficiency, and an optimization model for scheduling the industrial wireless network system is obtained. The optimization model carries each optimization parameter. The dynamic blockchain mechanism is used to dynamically adjust the leader edge server and the blockchain timeout window according to the network state and task demand, which not only guarantees safety and trustworthiness, but also improves the flexibility and adaptability of the system. Based on a preset multi-agent Markov decision process model, the optimization model is restructured to obtain a target optimization model. Based on the observation information of each industrial device in the industrial wireless network system collected in real time, a preset round-robin multi-agent deep reinforcement learning algorithm model is used to optimize the target optimization model, and target parameter values corresponding to each optimization parameter for scheduling the industrial wireless network system are obtained, thereby improving the trusted processing efficiency of the tasks in the industrial wireless network system.
[0147] The above embodiments are only exemplary embodiments of the present application and are not intended to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the present application within the spirit and protection scope of the present application, and such modifications or equivalent replacements should also be considered to fall within the protection scope of the present application.
Claims
1. A trusted scheduling method for industrial wireless networks based on dynamic blockchain, characterized in that, The method includes: Construct an industrial wireless network system based on a dynamic blockchain mechanism; Based on the preset constraints corresponding to the task and resource joint scheduling phase of the industrial wireless network system and the consensus phase of the dynamic blockchain mechanism, a model is constructed with the goal of maximizing the reliable processing efficiency of tasks, resulting in an optimized model for scheduling the industrial wireless network system. The optimized model carries various parameters to be optimized. The optimization model is reconstructed based on a preset multi-agent Markov decision process model to obtain the target optimization model. Based on the observation information of each industrial device in the industrial wireless network system collected in real time, the target optimization model is optimized by a preset round-robin multi-agent deep reinforcement learning algorithm model to obtain the target parameter values corresponding to each parameter to be optimized for scheduling the industrial wireless network system.
2. The method as described in claim 1, characterized in that, Before constructing the model based on the preset constraints corresponding to the task and resource joint scheduling phase of the industrial wireless network system and the consensus phase of the dynamic blockchain mechanism, with the goal of maximizing the reliable processing efficiency of tasks, the method further includes: constructing each preset constraint. The construction of each preset constraint condition specifically includes: Constraints are constructed based on the proportion of tasks offloaded from each industrial device to each edge server, resulting in the task offload ratio constraint conditions for the joint scheduling phase of tasks and resources. Based on the bandwidth ratio allocated by each edge server to each industrial device, the bandwidth allocation ratio constraint conditions are constructed to obtain the bandwidth allocation ratio constraint conditions for the task and resource joint scheduling phase. Constraints are constructed based on the predetermined maximum computing frequency parameters corresponding to each of the industrial devices, resulting in local computing frequency constraints corresponding to each of the industrial devices. Constraints are constructed based on the computing energy consumption, task transmission energy consumption and maximum battery capacity of each of the industrial devices, resulting in energy consumption constraints corresponding to each of the industrial devices. Based on the predetermined successful processing task counting function, the preset Byzantine fault tolerance coefficient, the identification parameter of each edge server completing the unloading task of each industrial equipment, and the task ratio of each industrial equipment unloading to each edge server, the blockchain security constraints of the consensus stage are constructed. Constraints are constructed based on the maximum computing frequency of each edge server, the binary indicator variable indicating whether each edge server acts as the leader, and the computing frequency allocated to the block when each edge server acts as the leader, to obtain computing frequency allocation constraints corresponding to each edge server. Constraints are constructed based on the predetermined task deadlines of each of the industrial devices to obtain the task deadline constraints corresponding to each of the industrial devices. Constraints are constructed based on the trust score and predetermined threshold of the number of successful task completions corresponding to each edge server, resulting in trustworthiness constraints corresponding to each edge server.
3. The method as described in claim 1, characterized in that, Before constructing a model based on the preset constraints corresponding to the task and resource joint scheduling phase of the industrial wireless network system and the consensus phase of the dynamic blockchain mechanism, with the goal of maximizing the trustworthy processing efficiency of tasks, the method further includes: constructing a trustworthy processing efficiency function for tasks. The reliability processing efficiency function for the construction task specifically includes: The trusted computing process function is obtained by constructing a function based on a preset trustworthiness verification delay function, a preset task unloading and transmission delay function, and a preset edge computing delay function. Based on the trusted computing process function, the preset transaction record report delay function, the preset maximum allowed consensus waiting time function, and the preset maximum actual consensus waiting time function, a function is constructed to obtain the actual consensus waiting delay function; Based on the actual consensus waiting delay function and the preset block generation delay function, a function is constructed to obtain the edge trusted computing delay function; Based on the edge trusted computing delay function and the local delay function corresponding to each industrial device, a function is constructed to obtain the task trusted computing delay function corresponding to each industrial device; Based on the task's trusted computation delay function and the task's data volume corresponding to the same industrial equipment, a function is constructed to obtain the trusted processing efficiency function corresponding to the same industrial equipment.
4. The method as described in claim 1, characterized in that, Before reconstructing the optimization model based on a preset multi-agent Markov decision process model, the method further includes: constructing a preset multi-agent Markov decision process model; The construction of the pre-defined multi-agent Markov decision process model specifically includes: Construct the set of agents, the set of agent observation states, and the set of agent actions of the preset multi-agent Markov decision process model; The reward function of the preset multi-agent Markov decision process model is obtained by constructing a function based on a preset trusted computing reward function, a preset timeout penalty function, and a preset consensus penalty function. The preset multi-agent Markov decision process model is obtained by constructing a model based on the set of agents, the set of agent observation states, the set of agent actions, and the reward function.
5. The method as described in claim 2, characterized in that, Before optimizing the target optimization model using a preset rotating multi-agent deep reinforcement learning algorithm model based on the observation information of each industrial device in the industrial wireless network system collected in real time, the method further includes: constructing a preset rotating multi-agent deep reinforcement learning algorithm model. The construction of the preset rotating multi-agent deep reinforcement learning algorithm model specifically includes: Construct an initial model corresponding to each of the aforementioned edge servers, the initial model including an initial actor neural network, two initial critic neural networks, and two initial target critic neural networks; The initial model corresponding to the predetermined initial leader agent in each of the edge servers is initialized to obtain the initial model parameters corresponding to the initial leader agent. Based on historical experience data, initial model parameters, preset first loss function, preset second loss function, preset third loss function, and preset entropy regularization loss function, the initial models are trained to obtain target deep reinforcement learning models corresponding to the edge servers, thereby obtaining the preset rotating multi-agent deep reinforcement learning algorithm. The target deep reinforcement learning model includes: a target actor neural network, each critic neural network, and a target critic neural network corresponding to each of the critic neural networks for stabilizing the critic neural networks.
6. The method as described in claim 5, characterized in that, The process of training each initial model based on historical experience data, initial model parameters, a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function to obtain a target deep reinforcement learning model corresponding to each edge server, thereby obtaining the preset rotating multi-agent deep reinforcement learning algorithm model, specifically includes: In the first decision-making period, a leader election is conducted based on the initial computing frequency, initial trust score, and initial channel state of each edge server to obtain the initial leader agent; The initial model corresponding to the initial leader agent is trained based on the preset first loss function, the preset second loss function, the preset third loss function, and the preset entropy regularization loss function to obtain the current first model corresponding to the initial leader agent; The model parameters of the current first model are sent to each first edge server that is not the initial leader agent. Based on the model parameters, the initial model corresponding to each first edge server is updated to obtain the current first model corresponding to each first edge server. The model parameters include the first model parameters, the second model parameters, the third model parameters, and the first entropy regularization coefficient corresponding to the initial leader agent. During periods other than the first decision-making phase, a new leader election is conducted based on the current computing frequency, current trust score, and current channel state of each edge server to obtain the current leader agent. The current first model corresponding to the current leader agent is then trained using a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function to update the current first model. The updated model parameters of the current first model are then distributed to each of the second edge servers that are not the current leader agent. Based on these model parameters, the current first model corresponding to each of the second edge servers is updated. This process is iterated until the algorithm converges, training the preset rotating multi-agent deep reinforcement learning algorithm model.
7. The method as described in claim 6, characterized in that, The target optimization model is optimized using a pre-defined round-robin multi-agent deep reinforcement learning algorithm model based on the observation information of each industrial device in the industrial wireless network system acquired in real time. This optimization yields target parameter values corresponding to each parameter to be optimized, used for scheduling the industrial wireless network system. Specifically, this includes: Based on the observation information of each industrial device in the industrial wireless network system collected in real time, the target actor neural network in the preset round-robin multi-agent deep reinforcement learning algorithm model is used to predict the state, and obtain the target parameter values corresponding to each of the parameters to be optimized for scheduling the industrial wireless network system. The parameters to be optimized include: task partitioning ratio, communication bandwidth allocation ratio, computation frequency of edge servers for offloading task processing and block generation, dynamic consensus waiting window in the blockchain, and leader agent elected by the edge servers.
8. A trusted scheduling device for industrial wireless networks based on dynamic blockchain, characterized in that, include: An industrial wireless network system building module is used to build an industrial wireless network system based on a dynamic blockchain mechanism. The model building module is used to build a model based on the preset constraints corresponding to the task and resource joint scheduling phase of the industrial wireless network system and the consensus phase of the dynamic blockchain mechanism, with the goal of maximizing the reliable processing efficiency of tasks, and to obtain an optimized model for scheduling the industrial wireless network system. The optimized model carries various parameters to be optimized. The model reconstruction module is used to reconstruct the optimization model based on a preset multi-agent Markov decision process model to obtain the target optimization model. The optimization module is used to optimize the target optimization model based on the observation information of each industrial device in the industrial wireless network system collected in real time, using a preset round-robin multi-agent deep reinforcement learning algorithm model, to obtain target parameter values corresponding to each parameter to be optimized for scheduling the industrial wireless network system.
9. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the trusted scheduling method for industrial wireless networks based on dynamic blockchain as described in any one of claims 1-7.
10. An electronic device, characterized in that, It includes at least a memory and a processor, wherein the memory stores a computer program, and the processor, when executing the computer program in the memory, implements the steps of the trusted scheduling method for industrial wireless networks based on dynamic blockchain as described in any one of claims 1-7.
Citation Information
Patent Citations
Heterogeneous task and resource end edge collaborative scheduling method based on digital twinning
CN116156563A
Block chain stable fragmentation method based on deep reinforcement learning and reputation mechanism
CN116506444A
Industrial internet data sharing optimization method based on intelligent fragmentation decision block chain
CN117478697A
Industrial wireless network trusted scheduling method and device based on dynamic block chain
CN119211957A
Adaptive-learning intelligent scheduling unified computing frame and system for industrial personalized customized production
US20220413455A1
Cited By
Sea mobile edge calculation dynamic unloading method for accidental tasks
CN121722580A
Space-time enhanced multi-agent reinforcement learning-based empty box allocation scheme optimization method
CN121937051A