A Digital Twin Network Trusted Scheduling Method and Device
The digital twin network scheduling method employs a distributed blockchain and multi-agent deep reinforcement learning to optimize task and resource allocation, addressing inefficiencies and trust issues, thereby enhancing network performance and reducing costs.
Patent Information
- Application Number
- CN202410951800.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-07-16
AI Technical Summary
There are problems in the digital twin network with low efficiency and high cost of task scheduling. Especially in dynamic and open scenarios, the trust between edge servers and terminal devices cannot be guaranteed, and the task processing delay and blockchain throughput cannot be effectively optimized.
Build a digital twin network system, adopt the diffusion blockchain mechanism to enable edge servers and terminal devices to join the blockchain, combine the multi-agent Markov decision-making process and deep reinforcement learning model, optimize task scheduling and resource allocation, and use lightweight diffusion to entrust the Byzantine fault-tolerant blockchain solution to reduce switching authentication overhead and optimize the blockchain consensus process.
It realizes the minimization of task trusted processing latency and maximizes blockchain transaction throughput, reduces communication overhead, and improves the overall performance of digital twin networks.
Smart Images

Figure CN119109920B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network scheduling, and in particular, to a trusted scheduling method, device, storage medium, and electronic device for a digital twin network. Background Art
[0002] Benefiting from the progress of new generation information technologies such as 5G, big data, and artificial intelligence, the implementation of digital twins has gradually become possible and is widely used in scenarios such as edge-cloud collaborative computing. However, most terminal devices are still restricted by manufacturing costs and internal space, and their computing and caching resources are still too limited to handle a large number of complex tasks, unable to fully meet the growing needs of people. To address this challenge, deploying edge servers with wide communication coverage, high computing capabilities, and large-capacity caching capabilities to establish a digital twin network for collaborative processing of computing tasks generated by terminal devices has become a common existing solution.
[0003] In addition, the dynamicity and openness of the digital twin network make it impossible to guarantee the trust between edge servers and terminal devices. During the task scheduling process, the task scheduling delay is not fully considered, resulting in low task processing efficiency and high costs. Summary of the Invention
[0004] In view of this, the present invention provides a trusted scheduling method, device, storage medium, and electronic device for a digital twin network, mainly aiming to solve the problems of low efficiency and high cost in the current digital twin network system during the network scheduling process.
[0005] To solve the above problems, the present application provides a trusted scheduling method for a digital twin network, including:
[0006] Constructing a digital twin network system;
[0007] Based on the digital twin network system, establishing a diffusion blockchain mechanism to enable each edge server and terminal device to join the blockchain to ensure the security and trustworthiness of network scheduling;
[0008] Based on preset constraint conditions corresponding to the task scheduling phase of the digital twin network system, the resource scheduling phase of the digital twin network system, and the blockchain consensus phase, the preset total delay function of the digital twin network system, and the preset blockchain throughput function, constructing a model to obtain an optimization model for digital twin network scheduling, where the optimization model carries each parameter to be optimized;
[0009] Based on a preset multi-agent Markov decision process model, reconstructing the optimization model to obtain a target optimization model for digital twin network scheduling;
[0010] Based on the observed status information of each terminal device in the digital twin network system collected in real time, a preset multi-agent deep reinforcement learning model is used to optimize the target optimization model to obtain target parameter values corresponding to each of the parameters to be optimized.
[0011] To solve the above problems, the present application provides a digital twin network trusted scheduling device, including:
[0012] A system construction module, configured to construct a digital twin network system;
[0013] A diffusion blockchain mechanism construction module, configured to establish a diffusion blockchain mechanism based on the digital twin network system, so that each edge server and terminal device join the blockchain to ensure the security and trustworthiness of network scheduling;
[0014] A model construction module, configured to perform model construction based on preset constraint conditions corresponding to the task scheduling phase of the digital twin network system, the resource scheduling phase of the digital twin network system, and the blockchain consensus phase, a preset total delay function of the digital twin network system, and a preset blockchain throughput function, to obtain an optimization model for digital twin network scheduling, where the optimization model carries each parameter to be optimized;
[0015] A model reconstruction module, configured to reconstruct the optimization model based on a preset multi-agent Markov decision process model to obtain a target optimization model for digital twin network scheduling;
[0016] An optimization module, configured to optimize the target optimization model based on the observed status information of each terminal device in the digital twin network system collected in real time by using a preset multi-agent deep reinforcement learning model to obtain target parameter values corresponding to each of the parameters to be optimized.
[0017] To solve the above problems, the present application provides a storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above digital twin network trusted scheduling method are implemented.
[0018] To solve the above problems, the present application provides an electronic device including at least a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program stored on the memory, the steps of the above digital twin network trusted scheduling method are implemented.
[0019] Beneficial effects in this application: This application addresses the issues of end-edge collaborative scheduling of tasks and resources and blockchain performance optimization in highly dynamic and open digital twin network scenarios. It fully considers constraints such as task migration, communication bandwidth, computing frequency, cache capacity, task deadline, and blockchain operation latency, formulates a joint optimization problem, and jointly optimizes the task migration ratio, communication bandwidth allocation, computing frequency allocation, cache decision, block size, and block generation interval to minimize the trusted processing latency of tasks and maximize the blockchain transaction throughput. For traditional high-overhead blockchain solutions, a digital twin-assisted lightweight diffusion delegated Byzantine fault-tolerant blockchain solution is proposed. By migrating the identity authentication process to the digital twin layer, the handover authentication overhead caused by mobility is reduced, and at the same time, the blockchain consensus process is optimized to reduce the communication overhead in the consensus process. For the problems of difficult modeling and algorithm state space explosion caused by the complex coupling of multi-dimensional resources and blockchain parameters, it is transformed into a multi-agent Markov decision process problem, and a multi-agent deep reinforcement learning model is proposed to seek the optimal scheduling strategy to improve the overall performance of the digital twin network.
[0020] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Brief Description of the Drawings
[0021] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0022] Figure 1 It shows a schematic flow chart of a trusted scheduling method for a digital twin network provided by an embodiment of the present application;
[0023] Figure 2 It shows a schematic scenario diagram of a digital twin network system for digital twin-assisted edge computing based on blockchain provided by an embodiment of the present application;
[0024] Figure 3 It shows a schematic diagram of a lightweight blockchain of diffusion delegated Byzantine fault tolerance provided by an embodiment of the present application;
[0025] Figure 4 It shows a schematic diagram of the algorithm structure of the dual actor-critic algorithm based on maximum entropy provided by an embodiment of the present application;
[0026] Figure 5The block diagram of a digital twin network trusted scheduling method device according to another embodiment of the present application is shown. Detailed implementation manners
[0027] Various solutions and features of the present application are described herein with reference to the accompanying drawings.
[0028] It should be understood that various modifications can be made to the embodiments applied herein. Therefore, the above description should not be regarded as a limitation, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of the present application.
[0029] The accompanying drawings included in the specification and constituting a part of the specification show the embodiments of the present application, and together with the general description of the present application given above and the detailed description of the embodiments given below are used to explain the principles of the present application.
[0030] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting examples with reference to the accompanying drawings.
[0031] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application.
[0032] When combined with the accompanying drawings, the above and other aspects, features, and advantages of the present application will become more apparent in view of the following detailed description.
[0033] Hereinafter, specific embodiments of the present application are described with reference to the accompanying drawings; however, it should be understood that the embodiments applied are only examples of the present application, and it can be implemented in various ways. Well-known and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details applied herein are not intended to be limiting, but only as a basis and representative basis for the claims to teach those skilled in the art to use the present application in substantially any suitable detailed structure in a variety of ways.
[0034] This specification may use the phrase "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", which may all refer to one or more of the same or different embodiments according to the present application.
[0035] The embodiments of the present application provide a digital twin network trusted scheduling method, as Figure 1 shown, including:
[0036] Step S101: Construct a digital twin network system;
[0037] In the specific implementation process of this step, the digital twin network system includes multiple edge servers and multiple terminal devices. The edge servers are powered by the power grid and are connected to each other by wire. The multiple edge servers include an edge computing server, a base station, and a blockchain node server. The edge computing server provides edge computing and content caching services. The base station server provides communication services. The blockchain node server provides security guarantees for communication, computing, and caching services. The terminal device generates tasks and performs local computing on the tasks, and supports migrating the tasks to the edge server through a wireless channel for edge computing. The digital twin is deployed on the edge server of the digital twin network system and is characterized as a virtualization model established for the edge servers and terminal devices in the network, and is used to evaluate the operating state and resource state of the digital twin network system. It can support the training of deep reinforcement learning methods to carry out end-edge collaborative scheduling of resources and tasks. At the same time, it can be used to execute the consensus process between blockchain nodes.
[0038] Step S102: Based on the digital twin network system, establish a diffusion blockchain mechanism so that each edge server and each terminal device join the blockchain to ensure the security and credibility of network scheduling.
[0039] In the specific implementation process of this step, in the blockchain consensus stage, each edge server and each terminal device are added to the blockchain system, so that the edge computing process is based on the diffusion blockchain mechanism to achieve the security and credibility of network scheduling.
[0040] Step S103: Based on the preset constraint conditions corresponding to the task scheduling stage of the digital twin network system, the resource scheduling stage of the digital twin network system, and the blockchain consensus stage, the preset total delay function of the digital twin network system, and the preset blockchain throughput function, perform model construction to obtain an optimization model for digital twin network scheduling. The optimization model carries each parameter to be optimized.
[0041] In the specific implementation process of this step, the blockchain consensus stage includes a migration and computing stage, a pre-preparation stage, a preparation stage, a holding stage, and a block generation interval stage. The preset constraint conditions include: a task migration ratio constraint condition, a cache constraint condition, a bandwidth ratio constraint condition, a first computing resource constraint, a second computing resource constraint condition, a deadline constraint condition, and a blockchain stability constraint condition; the online transaction stability of the blockchain is fully considered. Specifically, based on the preset total delay function and the preset blockchain throughput function, an optimization model for digital twin network scheduling is constructed with the goal of minimizing the task trusted processing delay and maximizing the blockchain throughput. The optimization model for digital twin network scheduling carries each parameter to be optimized. Each of the parameters to be optimized includes: a task migration ratio, a bandwidth allocation ratio, a computing resource allocation ratio, a cache decision, and a blockchain variable decision. This application adopts a diffusion Byzantine fault-tolerant blockchain construction method, which reduces the handover authentication overhead caused by the movement of terminal devices by migrating the identity authentication process to the digital twin layer, and at the same time optimizes the blockchain consensus process and reduces the communication overhead in the consensus process.
[0042] Step S104: Based on a preset multi-agent Markov decision process model, reconstruct the optimization model to obtain an objective optimization model for digital twin network scheduling;
[0043] In the specific implementation process of this step, based on the preset security parameters and the stability constraint conditions, a function is constructed to obtain a blockchain stability penalty function; based on the cache decision parameters estimated by content popularity, the data volume parameter of the cached content, the maximum cache capacity parameter of the edge server, the storage resource estimation deviation parameter, the content cache target decision parameter, and the cache constraint conditions, a function is constructed to obtain a cache penalty function; based on the blockchain stability penalty function and the cache penalty function, a function is constructed to obtain an objective penalty function; based on the preset task trusted processing delay weight coefficient, the blockchain transaction throughput weight coefficient, the preset total delay function, and the preset blockchain throughput function, a function is constructed to obtain a reward function; based on the reward function, the agent set, the system state set, the action set, and the state transition probability of the digital twin network system, a model is constructed to obtain the preset multi-agent Markov decision process model; based on the preset multi-agent Markov decision process model and the reward function, the optimization model of the digital twin network scheduling is transformed to obtain an objective optimization model of the digital twin network scheduling.
[0044] Step S105: Based on the observed state information of each terminal device in the digital twin network system collected in real time, use a preset multi-agent deep reinforcement learning model to optimize the objective optimization model to obtain the objective parameter values corresponding to each parameter to be optimized.
[0045] In the specific implementation of this step, the observed state information of each terminal device in the digital twin network system collected in real time is used by a preset multi-agent deep reinforcement learning model for state prediction to obtain each target parameter value that meets the observed state information; based on each target parameter value, the optimization model of the target digital twin network scheduling is used for calculation to obtain a target reward value; the target parameter value is executed to obtain the target observed state information of each terminal device in the next state; each observed state information, each data target parameter value, the target reward value, and the target observed state information are entered into a preset experience database; this lays a foundation for the offline update of the preset multi-agent deep reinforcement learning model in specific applications in the future.
[0046] This application addresses the problems of end-edge collaborative scheduling of tasks and resources and blockchain performance optimization in highly dynamic and open digital twin network scenarios. It fully considers constraints such as task migration, communication bandwidth, computing frequency, cache, task deadline, task processing delay, and blockchain stability, formulates a joint optimization problem, and jointly optimizes the task migration ratio, communication bandwidth allocation, computing frequency allocation, cache decision, block size, and block generation interval to minimize the trusted processing delay of tasks and maximize the blockchain transaction throughput. For traditional high-overhead blockchain solutions, a digital twin-assisted lightweight diffusion delegated Byzantine fault tolerance blockchain solution is proposed. By migrating the identity authentication process to the digital twin layer, the handover authentication overhead caused by mobility is reduced, and at the same time, the blockchain consensus process is optimized to reduce the communication overhead in the consensus process. For the problems of difficult modeling and algorithm state space explosion caused by the complex coupling of multi-dimensional resources and blockchain parameters, it is transformed into a multi-agent Markov decision process problem, and a deep reinforcement learning algorithm is proposed to seek the optimal scheduling strategy to improve the overall performance of the digital twin network.
[0047] Another embodiment of this application provides another digital twin network trusted scheduling method, including:
[0048] Step S201: Construct a digital twin network system;
[0049] In the specific implementation of this step, as Figure 2The figure shows a scenario diagram of the digital twin network system of blockchain-based edge computing assisted by digital twins of the present application. The digital twin network system includes multiple edge servers and multiple terminal devices. The edge servers are powered by the power grid and are connected by wire to each other. The multiple edge servers include edge computing servers, base stations, and blockchain node servers. The edge computing servers provide edge computing and content caching services. The base station servers provide communication services. The blockchain node servers provide security guarantees for communication, computing, and caching services. The terminal devices generate tasks and perform local computing on the tasks, and support migrating the tasks to the edge servers through wireless channels for edge computing. The digital twin is deployed on the edge servers and is characterized as a virtualization model established for the edge servers and terminal devices in the network, and is used to evaluate their operating status and resource status. It can support the training of deep reinforcement learning methods to carry out end-edge collaborative scheduling of the network. At the same time, it can be used to execute the consensus process between blockchain nodes.
[0050] Step S202: Construct preset constraint conditions corresponding to the task scheduling phase of the digital twin network system, the resource scheduling phase of the digital twin network system, and the blockchain consensus phase respectively.
[0051] In the specific implementation process of this step, the blockchain consensus phase includes a migration and computing phase, a pre-preparation phase, a preparation phase, a holding phase, and a block generation interval phase. As Figure 3 The figure shows a schematic diagram of a lightweight blockchain for diffusion delegated Byzantine fault tolerance of the present application. The preset constraint conditions include: a task migration ratio constraint condition, a cache constraint condition, a bandwidth ratio constraint condition, a first computing resource constraint, a second computing resource constraint condition, a deadline constraint condition, and a blockchain stability constraint condition. Specifically, based on the task ratio of each terminal device migrating to each edge server, a constraint condition is constructed to obtain the task migration ratio constraint condition. The mathematical expression of the task migration ratio constraint condition can be shown by the following formula (1):
[0052]
[0053] where u m,n represents the task ratio migrated from terminal device m to edge server n. When n = 0, it means that this part of the task is locally computed.
[0054] Based on the cache decision parameters estimated by content popularity, the data volume parameters of the cached content, the maximum cache capacity parameters of the edge servers, the storage resource estimation deviation parameters, and the content cache target decision parameters, a constraint condition is constructed to obtain the cache constraint condition. The mathematical expression of the cache constraint condition can be shown by the following formula (2):
[0055]
[0056] Among them, w k,m is the cache decision parameter for content popularity estimation; when w k,m = 1, content k is popular enough for terminal device m, and it is recommended to be cached by edge server n; the content caching target decision parameter for content k by the final algorithm is v n,k ; is the data volume parameter of the cached content of cached content k, S n is the maximum cache capacity parameter of edge server n, ΔS n is the storage resource estimation deviation parameter, that is, the size of the cached content cannot exceed the maximum storage capacity limit.
[0057] Based on the bandwidth ratio parameter allocated by the edge server to each of the terminal devices, a constraint condition is constructed to obtain a bandwidth ratio constraint condition; the mathematical expression of the bandwidth ratio constraint condition can be shown by the following formula (3):
[0058]
[0059] Among them, b m,n is the bandwidth ratio parameter allocated by edge server n to terminal device m.
[0060] Based on the edge computing resource parameter estimated in the preset digital twin layer that the edge server allocates to each terminal device, the computing resource estimation deviation parameter of the digital twin layer, and the maximum computing frequency parameter of the edge server, a constraint condition is constructed to obtain a first computing resource constraint condition and a second computing resource constraint condition; the mathematical expression of the first computing resource constraint condition can be shown by the following formula (4):
[0061]
[0062] The mathematical expression of the second computing resource constraint condition can be shown by the following formula (5):
[0063]
[0064] Among them, f m,n represents the edge computing resource parameter estimated in the digital twin layer that edge server n allocates to terminal device m; Δf m,n represents the computing resource estimation deviation parameter of the digital twin layer; represents the maximum computing frequency parameter of edge server n.
[0065] Based on the deadline parameter of the tasks executed by each of the terminal devices, a constraint condition is constructed to obtain a task deadline constraint condition; the mathematical expression of the task deadline constraint condition can be shown by the following formula (6):
[0066]
[0067] Among them, represents the deadline parameter of the task executed by the terminal device m, that is, the longest task processing delay that the terminal device m can accept;
[0068] Based on the blockchain operation delay parameter and the block generation interval parameter, a constraint condition is constructed to obtain the blockchain stability constraint condition. The mathematical expression of the stability constraint condition can be shown by the following formula (7):
[0069] C7: T LTF ≤ωT BI (7);
[0070] Among them, T LTF is the blockchain operation delay parameter, T BI is the block generation interval parameter, ω is an integer, representing several intervals, that is, a block should be published and verified within multiple consecutive block generation intervals.
[0071] The blockchain consists of full nodes served by edge servers and lightweight nodes served by terminal devices: The full nodes are responsible for recording transaction information during communication, computing, and caching of blocks, and participating in the decentralized consensus process; The lightweight nodes do not participate in the blockchain consensus process, but can use the blockchain to verify the credibility of edge servers when moving between different edge servers.
[0072] Step S203: Construct the preset total delay function of the digital twin network system;
[0073] In the specific implementation process of this step, a function is constructed based on the pre-preparation delay parameter, the preparation delay parameter, and the holding delay parameter to obtain the consensus delay function; The mathematical expression of the consensus delay function can be shown by the following formula (8):
[0074] T BC =T p-prep +T prep +T pers (8);
[0075] Among them, T p-prep is the pre-preparation delay parameter, T prep is the preparation delay parameter, T pers is the holding delay parameter;
[0076] The mathematical expression of the pre-preparation delay parameter can be shown by the following formula (9):
[0077]
[0078] Among them, the pre-preparation delay parameter T p-prep is related to the transaction data volume parameter D tr , the block data volume parameter D block , the computing frequency parameter C required for verification ver and the transmission rate parameter R between edge servers rsu .
[0079] The mathematical expression of the preparation delay parameter can be shown by the following formula (10):
[0080]
[0081] Among them, C ver represents the computing frequency required for validity verification
[0082] The mathematical expression of the holding delay can be shown by the following formula (11):
[0083]
[0084] Based on the consensus delay function and the block generation interval delay parameter, a function is constructed to obtain the blockchain operation delay function; the mathematical expression of the blockchain operation delay function can be shown by the following formula (12):
[0085] T LTF = T BC + T BI (12);
[0086] Based on the preset result download delay function, the preset task computing delay function, and the preset task migration delay function, a function is constructed to obtain the edge computing delay function; the mathematical expression of the edge computing delay function can be shown by the following formula (13):
[0087]
[0088] The mathematical expression of the task migration delay function can be shown by the following formula (14):
[0089]
[0090] Among them, u m,n represents the migration ratio, represents the task size, and R m,n represents the transmission rate
[0091] The mathematical expression of the result download delay function can be shown by the following formula (15):
[0092]
[0093] Among them, represents the result data volume and R m,n represents the transmission rate.
[0094] The mathematical expression of the preset task calculation delay function can be shown by the following formula (16):
[0095]
[0096] Among them, is the edge computing delay deviation, is the estimated edge computing delay. C m represents the computing frequency required for the task.
[0097] Based on the estimated local computing delay parameter and the local computing delay deviation parameter, a local computing delay function is constructed; the mathematical expression of the local computing delay function can be shown by the following formula (17):
[0098]
[0099] Among them, is the estimated local computing delay, is the local computing delay deviation.
[0100] Based on the consensus delay function, the blockchain operation delay function, the edge computing delay function, and the local computing delay function, a preset total delay function is constructed. The mathematical expression of the preset total delay function can be shown by the following formula (18):
[0101]
[0102] Among them, is the edge computing delay of task m, which is jointly determined by the cache decision and the edge computing delays of all edge servers n. When edge server n decides to cache the content k requested by the terminal device, that is, v n,k = 1, only the result needs to be downloaded from edge server n. Otherwise, the delays of task migration, computing, and result download need to be considered simultaneously.
[0103] In the process of constructing the total delay function in this application, the blockchain consensus mechanism mainly includes the following steps:
[0104] Step 1: Pre - preparation stage. First, the terminal device offloads the task to the target edge server, and the edge server records the transaction information of task offloading, computing, and caching. Then, all transactions are uploaded to a randomly selected entrusted full node. For a transaction, the full node first verifies its signature. If valid, the full node will execute the smart contract, and all transactions will be packaged to generate a new block.
[0105] Step 2: Preparation stage. Using the generated new block, the entrusted full node broadcasts the signed block as a new block proposal to other major full nodes participating in edge computing for verification. The block proposal is transmitted together with the pre - preparation message, which includes the ID, signature of the entrusted full node, and the hash result of the new block. When the major full nodes receive the pre - preparation message and the new block, they first verify the signature and MAC of the block, then verify the block through the smart contract, and finally return the vote on the block proposal to the entrusted full node. If the new block proposal is accepted by most major full nodes, that is, the number of votes is not less than the threshold, the blockchain consensus will enter the next stage.
[0106] Step 3: Maintenance stage. After being voted by the major full nodes, the new block will be added to the chain for update. Then, the entrusted full node will spread this block to all other major full nodes, including the minor full nodes that do not participate in edge computing.
[0107] In the edge computing process described above, for the tasks generated by a single terminal device, it can be selected not to migrate, partially migrate, or fully migrate them to one or more edge servers for computing.
[0108] Step S204: Construct the preset blockchain throughput function;
[0109] In the specific implementation of this step, based on the transaction quantity parameter, block data volume parameter, and transaction data volume parameter, a function is constructed to obtain the new block generation quantity function; the mathematical expression of the new block generation quantity function can be shown by the following formula (19):
[0110]
[0111] where L all represents the number of transactions, which is represented by formula (20):
[0112]
[0113] Based on the new block generation quantity function, block data volume parameter, transaction data volume parameter, block generation interval parameter, and transmission rate parameter between preset edge servers, a function is constructed to obtain the preset blockchain throughput function. The mathematical expression of the preset blockchain throughput function can be shown by the following formula (21):
[0114]
[0115] Among them, represents that the blockchain throughput is determined by the number of transactions included in a unit block, T BI represents the block generation interval, L block represents the number of blocks generated, R rsu represents the transmission rate between edge servers.
[0116] Step S205: Based on the digital twin network system, establish a diffusion blockchain mechanism so that each edge server and each terminal device join the blockchain to ensure the security and trustworthiness of network scheduling;
[0117] In the specific implementation process of this step, in the blockchain consensus stage, each terminal device and each edge server are added to the blockchain system, so that the edge computing process is based on the diffusion blockchain mechanism to achieve the security and trustworthiness of in-network scheduling.
[0118] Step S206: Based on the preset constraint conditions corresponding to the task scheduling stage of the digital twin network system, the resource scheduling stage of the digital twin network system, and the blockchain consensus stage, the preset total delay function of the digital twin network system, and the preset blockchain throughput function, a model is constructed to obtain an optimization model for digital twin network scheduling. The optimization model carries each parameter to be optimized;
[0119] In the specific implementation process of this step, the mathematical expression of the optimization model for digital twin network scheduling can be shown by the following formula (22):
[0120]
[0121] represents minimizing the task trusted processing delay and maximizing the blockchain throughput. T represents the total trusted processing delay of the entire digital twin network task, H represents the throughput of the blockchain, is the set of parameters to be optimized, representing the task migration ratio, bandwidth allocation ratio, computing resource allocation ratio, cache decision, and blockchain variable decision respectively.
[0122] Step S207: Based on the preset security parameters and the stability constraint conditions, a function is constructed to obtain a blockchain stability penalty function;
[0123] In the specific implementation process of this step, the mathematical expression of the blockchain stability penalty function can be shown by the following formula (23):
[0124]
[0125] where ξ is a negative value given according to security requirements.
[0126] Step S208: Construct a function based on the cache decision parameters estimated according to content popularity, the data volume parameters of cached content, the maximum cache capacity parameters of the edge server, the storage resource estimation deviation parameters, the content cache target decision parameters, and the cache constraint conditions to obtain a cache penalty function;
[0127] In the specific implementation process of this step, the mathematical expression of the cache penalty function can be shown by the following formula (24):
[0128]
[0129] Step S209: Construct a function based on the blockchain stability penalty function and the cache penalty function to obtain a target penalty function;
[0130] In the specific implementation process of this step, the mathematical expression of the target penalty function can be shown by the following formula (25):
[0131]
[0132] Step S210: Construct a function based on the preset task trusted processing delay weight coefficient, the blockchain transaction throughput weight coefficient, the preset total delay function, and the preset blockchain throughput function to obtain a reward function;
[0133] In the specific implementation process of this step, the mathematical expression of the reward function can be shown by the following formula (26):
[0134]
[0135] where α and β are weight coefficients for balancing the task trusted processing delay and the blockchain transaction throughput according to specific tasks.
[0136] Step S211: Construct a model based on the reward function, the agent set, the system state set, the action set, and the state transition probability of the digital twin network system to obtain the preset multi-agent Markov decision process model;
[0137] In the specific implementation process of this step, the preset multi-agent Markov decision process model is defined by the tuple where is the set of agents, is a set of observed system states, is a set of actions, is the state transition probability, is the reward function. The set of agents is denoted as responsible for observing the state space and learning better action strategies to maximize the reward. The observation space includes data size, required computing resource frequency, task deadline, communication bandwidth, cache storage capacity and estimation deviation, computing resources and their estimation deviation, and signal-to-noise ratio. The set of observed system states can be represented by the following formula (27):
[0138]
[0139] where, C(t) = {C1,..., C M , C ver , C val}, representing the data sizes of cache content, tasks, and transactions; represents the task deadline, represents the maximum computing frequency of the terminal device; ΔF end (t) = {Δf1, Δf2,..., Δf M} and represent the computing resource estimation deviation of the digital twin. γ n (t) = {γ 1,n , γ 2,n ,..., γ M,n} represents the signal-to-noise ratio.
[0140] In addition, the mathematical expression of the total state space of all agents can be represented by the following formula (28):
[0141] O(t) = {O1(t), O2(t),..., O N (t)} (28);
[0142] The action space is the action executed by the agent at a certain moment, and the mathematical expression of the action space can be represented by the following formula (29):
[0143]
[0144] where, u n (t) = {u 1,n , u 2,n ,..., u M,n} represents the task migration ratio, b n (t) = {b 1,n , b 2,n ,..., b M,n} represents the bandwidth allocation ratio, f n (t) = {f 1,n , f 2,n ,..., f M,n} represents the computing resource allocation ratio, v n (t) = {v 1,n , v 2,n ,..., v K,n} is the caching decision. The action space set of all agents is defined as A(t) = {A1(t), A2(t),..., A N (t)}.
[0145] The state transition probability describes the probability that O(t) is converted to O(t + 1) when the current agent executes the operation A(t). The transition probability from the current state to the next state can be expressed by the following formula (30):
[0146] pr(O(t + 1)|O(t), A(t)) = ∫ O(t+1) f(O(t), A(t), s′)ds′ (30); where f() is the probability density function of state transition.
[0147] Step S212: Based on the preset multi-agent Markov decision process model and the reward function, perform model transformation on the optimization model of digital twin network scheduling to obtain the target optimization model of digital twin network scheduling;
[0148] In the specific implementation process of this step, the mathematical expression of the target optimization model of digital twin network scheduling can be shown by the following formula (31):
[0149]
[0150] It means to maximize the long-term cumulative reward under the condition of meeting the constraints, so as to minimize the task trustworthy processing delay and maximize the blockchain throughput.
[0151] Step S213: Construct a preset multi-agent deep reinforcement learning model;
[0152] In the specific implementation process of this step, such as Figure 4The figure shows a schematic diagram of the algorithm structure for training the preset multi-agent deep reinforcement learning model based on the maximum entropy dual actor-critic algorithm of the present application. A first initial critic network model, a second initial critic network model, and an initial actor network model are respectively constructed; based on historical experience data, a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function are calculated and processed to obtain each model parameter, and each of the model parameters includes a first model parameter corresponding to the first initial critic network model, a second model parameter corresponding to the second initial critic network model, a third model parameter corresponding to the initial actor network model, and a regularization coefficient corresponding to the preset entropy regularization loss function; specifically, based on historical experience data, the first loss function of the first initial critic network model is used for calculation and processing to obtain the first model parameter σ1 when the loss value of the first loss function is minimized; wherein, the mathematical expression of the first loss function can be shown by the following formula (32):
[0153]
[0154] where ρ is the discount factor, is the first state value function, is the first initial critic network model; is the expected value calculation function; the mathematical expression of the first state value function can be shown by the following formula (33):
[0155]
[0156] Based on historical experience data, the second loss function of the second initial critic network model is used for calculation and processing to obtain the second model parameter σ2 when the loss value of the second loss function is minimized; wherein, the mathematical expression of the second loss function can be shown by the following formula (34):
[0157]
[0158] where ρ is the discount factor, is the second state value function, and the mathematical expression of the second state value function can be shown by the following formula (35):
[0159]
[0160] Screening is performed based on the first model parameter and the second model parameter to obtain a target model parameter; specifically, the first model parameter and the second model parameter are compared, and the smaller model parameter is determined as the target model parameter. Based on the target model parameter and the historical experience data, calculation and processing are performed using the third loss function of the initial actor network model to obtain the third model parameter θ when the loss value of the third loss function is minimized; the mathematical expression of the third loss function can be shown by the following formula (36):
[0161]
[0162] Based on the historical experience data, calculation and processing are performed using a preset entropy regularization loss function to obtain a regularization coefficient when the loss value of the preset entropy regularization loss function is minimized. The mathematical expression of the preset entropy regularization loss function can be shown by the following formula (37):
[0163]
[0164] where is the policy entropy of maximum entropy reinforcement learning, and E0 is the target entropy;
[0165] Based on the first model parameter, the second model parameter, the third model parameter, and the regularization coefficient, model updates are performed to obtain the first current critic network model, the second current critic network model, the current actor model, and the current entropy regularization loss function; specifically, based on the first model parameter, the first initial critic network model is updated to obtain the first current critic network model; based on the second model parameter, the second initial critic network model is updated to obtain the second current critic network model; based on the third model parameter, the initial actor network model is updated to obtain the current actor model; based on the regularization coefficient, the preset entropy regularization loss function is updated to obtain the current entropy regularization loss function. Specifically, the following soft update function can be used to update each model, and the mathematical expression of the soft update function can be shown by the following formula (38):
[0166]
[0167] Iteratively update the first current critic network model, the second current critic network model, the current actor model, and the current entropy regularization loss function in a loop until the reward function value obtained by optimizing the target optimization model of the digital twin network scheduling with the state parameter values obtained by the updated current actor model for predicting the states of each of the historical experience data meets the preset conditions, and then determine the updated current actor network model as the target actor network model.
[0168] Step S214: Optimize the optimization model for the target digital twin network scheduling by using a preset multi-agent deep reinforcement learning model based on the observation status information of each terminal device in the digital twin network system collected in real time, so as to obtain target parameter values corresponding to each parameter to be optimized.
[0169] In the specific implementation process of this step, the observation status information of each terminal device in the digital twin network system collected in real time is used by a preset multi-agent deep reinforcement learning model for state prediction to obtain each target parameter value that meets the observation status information; based on each target parameter value, the optimization model for the target digital twin network scheduling is used for calculation to obtain a target reward value; execute the target parameter value to obtain the target observation status information of each terminal device in the next state; input each observation status information, each data target parameter value, the target reward value, and the target observation status information into a preset experience database; lay a foundation for the offline update of the preset multi-agent deep reinforcement learning model in specific applications in the future.
[0170] Another embodiment of the present application provides a digital twin network trusted scheduling device, as Figure 5 shown, including:
[0171] System construction module 1, used to construct a digital twin network system;
[0172] Diffusion blockchain mechanism construction module 2, used to establish a diffusion blockchain mechanism based on the digital twin network system, so that each edge server and each terminal device join the blockchain to ensure the security and trustworthiness of network scheduling;
[0173] Model construction module 3, used to construct a model based on preset constraint conditions corresponding to the task scheduling stage of the digital twin network system, the resource scheduling stage of the digital twin network system, and the blockchain consensus stage, the preset total delay function of the digital twin network system, and the preset blockchain throughput function, so as to obtain an optimization model for digital twin network scheduling, and each parameter to be optimized is carried in the optimization model;
[0174] Model reconstruction module 4, used to reconstruct the optimization model based on a preset multi-agent Markov decision process model to obtain a target optimization model for digital twin network scheduling;
[0175] Optimization module 5, used to optimize the target optimization model by using a preset multi-agent deep reinforcement learning model based on the observation status information of each terminal device in the digital twin network system collected in real time, so as to obtain target parameter values corresponding to each parameter to be optimized.
[0176] In the specific implementation process, the diffusion blockchain mechanism construction module 2 is specifically used for: in the blockchain consensus stage, adding each of the edge servers and each of the terminal devices to the blockchain system, so that the edge computing process is based on the diffusion blockchain mechanism to achieve secure and trustworthy network scheduling. In the specific implementation process, the device further includes: a constraint condition construction module, and the constraint condition construction module is specifically used for: constructing preset constraint conditions corresponding to the task scheduling stage of the digital twin network system, the resource scheduling stage of the digital twin network system, and the blockchain consensus stage respectively, specifically including: constructing constraint conditions based on the task ratio of each terminal device migrating to each edge server to obtain task migration ratio constraint conditions; constructing constraint conditions based on the cache decision parameters estimated by content popularity, the data volume parameters of cached content, the maximum cache capacity parameters of edge servers, the storage resource estimation deviation parameters, and the content cache target decision parameters to obtain cache constraint conditions; constructing constraint conditions based on the bandwidth ratio parameters allocated by edge servers to each of the terminal devices to obtain bandwidth allocation ratio constraint conditions; constructing constraint conditions based on the edge computing resource parameters allocated by edge servers to each of the terminal devices estimated in the preset digital twin layer, the computing resource estimation deviation parameters of the digital twin layer, the maximum computing frequency parameters of edge servers, the computing resource parameters allocated by the terminal devices themselves estimated in the digital twin layer, the computing resource estimation deviation parameters of the digital twin layer, and the maximum computing frequency parameters of the terminal devices to obtain a first computing resource constraint condition and a second computing resource constraint condition; constructing constraint conditions based on the deadline parameters of tasks executed by each of the terminal devices to obtain task deadline constraint conditions; constructing constraint conditions based on the blockchain operation delay parameters and the block generation interval parameters to obtain blockchain stability constraint conditions.
[0177] In the specific implementation process, the device further includes: a preset total delay function construction module, and the preset total delay function construction module is specifically used for: constructing a function based on the pre-preparation delay parameter, the preparation delay parameter, and the holding delay parameter to obtain a consensus delay function; constructing a function based on the consensus delay function and the block generation interval delay parameter to obtain a blockchain operation delay function; constructing a function based on a preset task calculation delay function, a preset task migration delay function, and a preset result download delay function to obtain an edge computing delay function; constructing a function based on the estimated local calculation delay parameter, the local calculation delay deviation parameter, and the local task calculation delay function to obtain a local calculation delay function; constructing a function based on the consensus delay function, the blockchain operation delay function, the edge computing delay function, and the local calculation delay function to obtain the preset total delay function.
[0178] In the specific implementation process, the device further includes: a preset blockchain throughput function construction module, and the preset blockchain throughput function construction module is specifically used for: constructing a function based on the transaction quantity parameter, the block data volume parameter, and the transaction data volume parameter to obtain a new block generation quantity function; constructing a function based on the new block generation quantity function, the block data volume parameter, the transaction data volume parameter, the block generation interval parameter, and the transmission rate parameter between preset edge servers to obtain the preset blockchain throughput function.
[0179] In the specific implementation process, the model construction module 3 is specifically used for: constructing a function based on the preset security parameter and the stability constraint condition to obtain a blockchain stability penalty function; constructing a function based on the cache decision parameter estimated by content popularity, the data volume parameter of the cached content, the maximum cache capacity parameter of the edge server, the storage resource estimation deviation parameter, the content cache target decision parameter, and the cache constraint condition to obtain a cache penalty function; constructing a function based on the blockchain stability penalty function and the cache penalty function to obtain a target penalty function; constructing a function based on the preset task trusted processing delay weight coefficient, the blockchain transaction throughput weight coefficient, the preset total delay function, and the preset blockchain throughput function to obtain a reward function; constructing a model based on the reward function, the agent set, the system state set, the action set, and the state transition probability of the digital twin network system to obtain the preset multi-agent Markov decision process model; performing model transformation on the optimization model of the digital twin network scheduling based on the preset multi-agent Markov decision process model and the reward function to obtain the target optimization model of the digital twin network scheduling.
[0180] In the specific implementation process, the device further includes: a preset multi-agent deep reinforcement learning model construction module, and the preset multi-agent deep reinforcement learning model construction module is specifically used for: respectively constructing a first initial critic network model, a second initial critic network model, and an initial actor network model; performing calculation processing based on historical experience data, a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function to obtain each model parameter, and each of the model parameters includes a first model parameter corresponding to the first initial critic network model, a second model parameter corresponding to the second initial critic network model, a third model parameter corresponding to the initial actor network model, and a regularization coefficient corresponding to the preset entropy regularization loss function; performing model update based on the first model parameter, the second model parameter, the third model parameter, and the regularization coefficient to obtain a first current critic network model, a second current critic network model, a current actor model, and a current entropy regularization loss function; iteratively updating the first current critic network model, the second current critic network model, the current actor model, and the current entropy regularization loss function in a loop until the reward function value obtained by optimizing the target optimization model with each state parameter value obtained by the updated current actor model predicting the states of each of the historical experience data meets a preset condition, and determining the updated current actor network model as the preset multi-agent deep reinforcement learning model.
[0181] In the specific implementation process, the preset multi-agent deep reinforcement learning model construction module is further used for: performing calculation processing based on the historical experience data using the first loss function of the first initial critic network model to obtain the first model parameter when the loss value of the first loss function is the smallest; performing calculation processing based on the historical experience data using the second loss function of the second initial critic network model to obtain the second model parameter when the loss value of the second loss function is the smallest; performing screening based on the first model parameter and the second model parameter to obtain a target model parameter; performing calculation processing based on the target model parameter and the historical experience data using the third loss function of the initial actor network model to obtain the third model parameter when the loss value of the third loss function is the smallest; performing calculation processing based on the historical experience data using the preset entropy regularization loss function to obtain the regularization coefficient when the loss value of the preset entropy regularization loss function is the smallest.
[0182] This application addresses the issues of end-edge collaborative scheduling of tasks and resources and blockchain performance optimization in highly dynamic and open digital twin network scenarios. It fully considers constraints such as task migration, communication bandwidth, computing frequency, cache capacity, task deadlines, and blockchain operation latency, formulates a joint optimization problem, and jointly optimizes the task migration ratio, communication bandwidth allocation, computing frequency allocation, cache decision-making, block size, and block generation interval to minimize the trusted processing latency of tasks and maximize the throughput of blockchain transactions. For traditional high-overhead blockchain solutions, a digital twin-assisted lightweight diffusion delegated Byzantine fault-tolerant blockchain solution is proposed. By migrating the identity authentication process to the digital twin layer, the handover authentication overhead caused by mobility is reduced, and at the same time, the blockchain consensus process is optimized to reduce the communication overhead during the consensus process. To address the challenges of difficult modeling and explosive algorithm state space caused by the complex coupling of multi-dimensional resources and blockchain parameters, it is transformed into a multi-agent Markov decision process problem, and a target actor network model is proposed to seek the optimal scheduling strategy to improve the overall performance of the digital twin network.
[0183] Another embodiment of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the following method steps:
[0184] Step 1: Construct a digital twin network system;
[0185] Step 2: Based on the digital twin network system, establish a diffusion blockchain mechanism to enable each edge server and each terminal device to join the blockchain to ensure the security and trustworthiness of network scheduling;
[0186] Step 3: Based on the preset constraint conditions corresponding to the task scheduling phase, the resource scheduling phase, and the blockchain consensus phase of the digital twin network system, the preset total latency function of the digital twin network system, and the preset blockchain throughput function, construct a model to obtain an optimization model for digital twin network scheduling, and the optimization model carries each parameter to be optimized;
[0187] Step 4: Based on a preset multi-agent Markov decision process model, reconstruct the optimization model to obtain a target optimization model for digital twin network scheduling;
[0188] Step 5: Based on the observed state information of each terminal device in the digital twin network system collected in real time, use a preset multi-agent deep reinforcement learning model to optimize the target optimization model to obtain target parameter values corresponding to each parameter to be optimized.
[0189] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0190] For the specific implementation process of the above method steps, reference can be made to the embodiments of any of the above digital twin network trusted scheduling methods, and this embodiment will not be repeated here.
[0191] This application aims at the problem of end-edge collaborative scheduling of tasks and resources and blockchain performance optimization in a highly dynamic and open digital twin network scenario. It fully considers constraints such as task migration, communication bandwidth, computing frequency, cache capacity, task deadline, and blockchain operation delay, formulates a joint optimization problem, and jointly optimizes the task migration ratio, communication bandwidth allocation, computing frequency allocation, cache decision-making, block size, and block generation interval to minimize the trusted processing delay of tasks and maximize the blockchain transaction throughput. For traditional high-overhead blockchain solutions, a digital twin-assisted lightweight diffusion delegated Byzantine fault-tolerant blockchain solution is proposed. By migrating the identity authentication process to the digital twin layer, the handover authentication overhead caused by mobility is reduced, and at the same time, the blockchain consensus process is optimized to reduce the communication overhead in the consensus process. Aiming at the problems of difficult modeling and algorithm state space explosion caused by the complex coupling of multi-dimensional resources and blockchain parameters, it is transformed into a multi-agent Markov decision process problem, and a target actor network model is proposed to seek the optimal scheduling strategy to improve the overall performance of the digital twin network.
[0192] Another embodiment of this application provides an electronic device, which at least includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program on the memory, the following method steps are implemented:
[0193] Step 1: Construct a digital twin network system;
[0194] Step 2: Based on the digital twin network system, establish a diffusion blockchain mechanism so that each edge server and each terminal device join the blockchain to ensure the security and trust of network scheduling;
[0195] Step 3: Based on the preset constraint conditions corresponding to the task scheduling stage of the digital twin network system, the resource scheduling stage of the digital twin network system, and the blockchain consensus stage, the preset total delay function of the digital twin network system, and the preset blockchain throughput function, a model is constructed to obtain an optimization model for digital twin network scheduling. The optimization model carries each parameter to be optimized;
[0196] Step 4: Based on the preset multi-agent Markov decision process model, perform model reconstruction on the optimization model to obtain an objective optimization model for digital twin network scheduling;
[0197] Step 5: Based on the observed state information of each terminal device in the digital twin network system collected in real time, use the preset multi-agent deep reinforcement learning model to optimize the objective optimization model to obtain the objective parameter values corresponding to each parameter to be optimized.
[0198] For the specific implementation process of the above method steps, reference can be made to the embodiments of any of the above digital twin network trusted scheduling methods, and this embodiment will not be repeated here.
[0199] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the present application within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.
Claims
1. A trusted scheduling method for a digital twin network, characterized in that, Including: Constructing a digital twin network system; Based on the digital twin network system, establishing a diffusion blockchain mechanism to enable each edge server and each terminal device to join the blockchain, so as to ensure the security and credibility of network scheduling; Based on the preset constraint conditions corresponding to the task scheduling stage of the digital twin network system, the resource scheduling stage of the digital twin network system, and the blockchain consensus stage respectively, the preset total delay function of the digital twin network system, and the preset blockchain throughput function, a model is constructed with the goal of minimizing the task trusted processing delay and maximizing the blockchain throughput, obtaining a multi-objective optimization model for digital twin network scheduling, and the optimization model carries each parameter to be optimized; Based on a preset multi-agent Markov decision process model, the optimization model is reconstructed to obtain a target optimization model for digital twin network scheduling; Based on the observed state information of each terminal device in the digital twin network system collected in real time, a preset multi-agent deep reinforcement learning model is used to optimize the target optimization model to obtain the target parameter values corresponding to each parameter to be optimized; Among them, the preset constraint conditions include: task migration ratio constraint condition, cache constraint condition, bandwidth allocation ratio constraint condition, first computing resource constraint condition, second computing resource constraint condition, and blockchain stability constraint condition; The parameters to be optimized include: task migration ratio, bandwidth allocation ratio, computing resource allocation ratio, cache decision, and blockchain variable decision; The mathematical expression of the multi-objective optimization model can be represented by the following formula: Among them, T represents the total delay of the trusted processing of the entire digital twin network task, and H represents the throughput of the blockchain. is the set of parameters to be optimized, representing the task migration ratio, bandwidth allocation ratio, computing resource allocation ratio, caching decision, and blockchain variable decision respectively. The mathematical expression of the total delay T can be represented by the following formula: Among them, is the edge computing delay of task m; is the local computing delay of task m; T LTF is the blockchain operation delay; The mathematical expression of the throughput of the blockchain can be represented by the following formula: Among them, indicates that the blockchain throughput is determined by the number of transactions included in a unit block, T BI indicates the block generation interval, L block indicates the number of blocks generated, R rsu indicates the transmission rate between edge servers, D block indicates the block data volume.
2. The method according to claim 1, wherein In the blockchain consensus stage, each edge server and each terminal device are added to the blockchain system, so that the edge computing process is based on the diffusion blockchain mechanism to achieve the security and credibility of network scheduling.
3. The method according to claim 1, wherein Before constructing the model based on the preset constraint conditions corresponding to the task scheduling stage of the digital twin network system, the resource scheduling stage of the digital twin network system, and the blockchain consensus stage respectively, the preset total delay function of the digital twin network system, and the preset blockchain throughput function, the method further includes: constructing the preset constraint conditions corresponding to the task scheduling stage of the digital twin network system, the resource scheduling stage of the digital twin network system, and the blockchain consensus stage respectively, specifically including: Based on the task ratio of each terminal device migrating to each edge server, a constraint condition is constructed to obtain a task migration ratio constraint condition; Based on the cache decision parameter estimated by content popularity, the data volume parameter of the cached content, the maximum cache capacity parameter of the edge server, the storage resource estimation deviation parameter, and the content cache target decision parameter, a constraint condition is constructed to obtain a cache constraint condition; Based on the bandwidth ratio parameter allocated by the edge server to each terminal device, a constraint condition is constructed to obtain a bandwidth allocation ratio constraint condition; Construct constraint conditions based on the edge computing resource parameters allocated by the edge server to each terminal device estimated in the preset digital twin layer, the calculation resource estimation deviation parameters of the digital twin layer, the maximum calculation frequency parameters of the edge server, the calculation resource parameters allocated by the terminal device itself estimated in the digital twin layer, the calculation resource estimation deviation parameters of the digital twin layer, and the maximum calculation frequency parameters of the terminal device, to obtain the first calculation resource constraint condition and the second calculation resource constraint condition; Construct constraint conditions based on the deadline parameters of the tasks executed by each of the terminal devices, to obtain the task deadline constraint condition; Construct constraint conditions based on the blockchain operation delay parameters and the block generation interval parameters, to obtain the blockchain stability constraint condition.
4. The method according to claim 1, wherein Before constructing a model based on the preset constraint conditions corresponding to the task scheduling stage of the digital twin network system, the resource scheduling stage of the digital twin network system, and the blockchain consensus stage, the preset total delay function of the digital twin network system, and the preset blockchain throughput function, the method further includes: constructing the preset total delay function of the digital twin network system and constructing the preset blockchain throughput function, specifically including: Construct a function based on the pre-preparation delay parameter, the preparation delay parameter, and the holding delay parameter, to obtain the consensus delay function; Construct a function based on the consensus delay function and the block generation interval delay parameter, to obtain the blockchain operation delay function; Construct a function based on the preset task calculation delay function, the preset task migration delay function, and the preset result download delay function, to obtain the edge computing delay function; Construct a function based on the estimated local calculation delay parameter, the local calculation delay deviation parameter, and the local task calculation delay function, to obtain the local calculation delay function; Construct a function based on the consensus delay function, the blockchain operation delay function, the edge computing delay function, and the local calculation delay function, to obtain the preset total delay function; Construct a function based on the transaction quantity parameter, the block data volume parameter, and the transaction data volume parameter, to obtain the new block generation quantity function; Construct a function based on the new block generation quantity function, the block data volume parameter, the transaction data volume parameter, the block generation interval parameter, and the transmission rate parameter between the preset edge servers, to obtain the preset blockchain throughput function.
5. The method according to claim 1, wherein The model reconstruction of the optimization model based on the preset multi-agent Markov decision process model to obtain the target optimization model for digital twin network scheduling specifically includes: Construct a function based on the preset security parameter and the stability constraint condition, to obtain the blockchain stability penalty function; Construct a function based on the cache decision parameter estimated by the content popularity, the data volume parameter of the cached content, the maximum cache capacity parameter of the edge server, the storage resource estimation deviation parameter, the content cache target decision parameter, and the cache constraint condition, to obtain the cache penalty function; Construct a function based on the blockchain stability penalty function and the cache penalty function, to obtain the target penalty function; Function construction is performed based on the preset task trustworthy processing delay weight coefficient, the blockchain transaction throughput weight coefficient, the preset total delay function, and the preset blockchain throughput function to obtain a reward function; Based on the reward function, the agent set, the system state set, the action set, and the state transition probability of the digital twin network system, model construction is carried out to obtain the preset multi-agent Markov decision process model; Based on the preset multi-agent Markov decision process model and the reward function, model transformation is performed on the optimization model of digital twin network scheduling to obtain the target optimization model of digital twin network scheduling.
6. The method according to claim 1, wherein Before optimizing the target optimization model using the preset multi-agent deep reinforcement learning model based on the observed state information of each terminal device in the digital twin network system collected in real time, the method further includes: constructing the preset multi-agent deep reinforcement learning model, specifically including: Constructing a first initial critic network model, a second initial critic network model, and an initial actor network model respectively; Based on historical experience data, a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function, calculation processing is carried out to obtain each model parameter. Each of the model parameters includes a first model parameter corresponding to the first initial critic network model, a second model parameter corresponding to the second initial critic network model, a third model parameter corresponding to the initial actor network model, and a regularization coefficient corresponding to the preset entropy regularization loss function; Based on the first model parameter, the second model parameter, the third model parameter, and the regularization coefficient, model update is carried out to obtain a first current critic network model, a second current critic network model, a current actor model, and a current entropy regularization loss function; The first current critic network model, the second current critic network model, the current actor model, and the current entropy regularization loss function are iteratively updated in a loop until the reward function value obtained by optimizing the target optimization model with the action parameter values obtained by the updated current actor model for predicting the states of each piece of the historical experience data meets a preset condition. At this time, the updated network models are determined as the preset multi-agent deep reinforcement learning model.
7. The method according to claim 6, wherein The calculation processing based on historical experience data, a preset first loss function, a preset second loss function, a preset third loss function, and a preset entropy regularization loss function to obtain each model parameter specifically includes: Based on historical experience data, calculation processing is carried out using the first loss function of the first initial critic network model to obtain the first model parameter when the loss value of the first loss function is the smallest; Based on historical experience data, calculation processing is carried out using the second loss function of the second initial critic network model to obtain the second model parameter when the loss value of the second loss function is the smallest; Based on the first model parameter and the second model parameter, screening is carried out to obtain the target model parameter; Based on the target model parameters and the historical experience data, perform calculation processing using the third loss function of the initial actor network model to obtain the third model parameters when the loss value of the third loss function is minimized. Based on the historical experience data, perform calculation processing using a preset entropy regularization loss function to obtain the regularization coefficient when the loss value of the preset entropy regularization loss function is minimized.
8. A network resource scheduling device, characterized in that It includes: A system construction module for constructing a digital twin network system. A diffusion blockchain mechanism construction module, based on the digital twin network system, to establish a diffusion blockchain mechanism, enabling each edge server and terminal device to join the blockchain to ensure the security and trustworthiness of network scheduling. A model construction module for constructing a model with the goal of minimizing the task trusted processing delay and maximizing the blockchain throughput based on preset constraint conditions corresponding to the task scheduling stage, the resource scheduling stage, and the blockchain consensus stage of the digital twin network system, the preset total delay function of the digital twin network system, and the preset blockchain throughput function, to obtain a multi-objective optimization model for digital twin network scheduling, where the optimization model carries various parameters to be optimized; the preset constraint conditions include: task migration ratio constraint condition, cache constraint condition, bandwidth allocation ratio constraint condition, first computing resource constraint condition, second computing resource constraint condition, and blockchain stability constraint condition; the parameters to be optimized include: task migration ratio, bandwidth allocation ratio, computing resource allocation ratio, cache decision, and blockchain variable decision; the mathematical expression of the multi-objective optimization model can be represented by the following formula: Among them, T represents the total delay of the entire trusted processing of the digital twin network task, and H represents the throughput of the blockchain. is the set of parameters to be optimized, representing the task migration ratio, bandwidth allocation ratio, computing resource allocation ratio, caching decision, and blockchain variable decision respectively. The mathematical expression of the total delay T can be represented by the following formula: Among them, is the edge computing delay of task m; is the local computing delay of task m; T LTF is the blockchain operation delay; The mathematical expression of the throughput of the blockchain can be represented by the following formula: Among them, represents that the blockchain throughput is determined by the number of transactions included in a unit block, T BI represents the block generation interval, L block represents the number of blocks generated, R rsu represents the transmission rate between edge servers, D block represents the block data volume; A model reconstruction module for reconstructing the optimization model based on a preset multi-agent Markov decision process model to obtain a target optimization model for digital twin network scheduling. An optimization module for optimizing the target optimization model using a preset multi-agent deep reinforcement learning model based on the observed state information of each terminal device in the digital twin network system collected in real time to obtain the target parameter values corresponding to each of the parameters to be optimized.
9. A storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the digital twin network trusted scheduling method according to any one of claims 1-7 above.
10. An electronic device, characterized in that, It at least includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program on the memory, it implements the steps of the digital twin network trusted scheduling method according to any one of claims 1-7 above.
Citation Information
Patent Citations
Heterogeneous task and resource end edge collaborative scheduling method based on digital twinning
CN116156563A
Block chain stable fragmentation method based on deep reinforcement learning and reputation mechanism
CN116506444A
Security service migration method and system in mobile edge computing scene
CN117750436A