Method for resource allocation of digital twin assisted end-to-end deterministic latency network slice in industrial internet of things

By optimizing resource allocation through digital twin technology and multi-agent deep reinforcement learning algorithms, the deterministic latency and reliability issues across network domains in the Industrial Internet of Things are solved, achieving efficient network resource management and maximizing system utility.

CN119031484BActive Publication Date: 2025-10-21CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410988743.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-10-21
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

Existing technologies are unable to meet the diverse needs of latency-sensitive services in the Industrial Internet of Things, especially in resource allocation across network domains, where it is difficult to guarantee deterministic latency and reliability requirements.

Method used

Digital twin technology is used to build an end-to-end network slicing architecture, combined with mobile edge computing. VNF perception of network request information is synchronized through digital twins, and resource allocation is optimized using random network calculations and multi-agent deep reinforcement learning algorithms to ensure deterministic latency while maximizing system utility.

Benefits of technology

It achieves the goal of improving system utility, optimizing network resource allocation, and meeting the latency and reliability requirements of diversified services in the Industrial Internet of Things while meeting deterministic latency requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119031484B_ABST
    Figure CN119031484B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of resource allocation methods of digital twin assisted end-to-end deterministic latency network slice in industrial internet of things, belong to mobile communication technical field.The method includes the following steps: S1: in the industrial internet of things scene, construct digital twin assisted end-to-end network slice architecture;S2: analysis the influence of the computing resource offset and synchronous delay of digital twin mapping process on the end-to-end delay of service;S3: introduce random network calculus SNC to characterize the upper bound of delay in service transmission;S4: with the goal of maximizing system utility while ensuring deterministic delay, build joint multi-dimensional resource allocation problem;S5: propose DT- PER-MASAC algorithm to solve optimization problem.The present application can realize the optimization of resource allocation strategy in end-to-end network slice, improve network utility while guaranteeing the certainty of service delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of mobile communication technology and relates to a resource allocation method for digital twin-assisted end-to-end deterministic delay network slicing in the industrial Internet of Things. Background Art

[0002] With the development of IoT technology, a massive number of IoT devices are connected to the network. Existing networks are unable to fully meet the stringent and diverse requirements of latency-sensitive services for latency, reliability, and other aspects. To meet the Quality of Service (QoS) requirements of transmission services, network service providers (NSPs) use network slicing technology to abstract the physical resources provided by infrastructure providers (InPs) into virtual resources. NSPs can configure virtual resources based on user needs and provide personalized services. As a key technology in 5G networks, network slicing typically requires resources from multiple network domains. There is a very complex trade-off between these resources and slice performance. Therefore, resource allocation is crucial to the resource utilization and network performance of network slices.

[0003] Due to the dynamic nature of business needs and device status information, it is currently difficult to ensure service QoS under information uncertainty. Meeting stringent and diverse business requirements has become a critical and urgent goal for network deployment. To achieve this goal, new technologies are needed to obtain real-time network information. The 6G era will be a digital era. Digital Twin (DT), one of the key technologies of 6G, will be integrated with communication networks to map entities in the physical network (including network elements, user devices, etc.) into a virtual space, enabling the construction of a Digital Twin Network (DTN) that is consistent with the physical network. In industrial IoT scenarios, the integration of IoT and DT, supported by Mobile Edge Computing (MEC), enables real-time monitoring of the status of network elements across the entire end-to-end network slice. Functional models provide perception data and decision-making, bringing more intelligent and efficient solutions to industrial applications.

[0004] The diversification of business demands places higher demands on network flexibility. Although existing technologies have explored aspects such as DT-assisted network slicing and resource allocation, the dynamic changes in IoT terminal service requests and service node status make it difficult to meet latency and reliability requirements for cross-domain resource allocation in end-to-end network slicing. Existing research only considers minimizing network latency or maximizing network resource utilization from a single network domain, which cannot meet the needs of new services with high latency and reliability requirements. A key issue in providing deterministic services is how to allocate the appropriate amount of resources to ensure QoS requirements within the network slice. Deterministic latency performance, rather than latency minimization, can meet the network performance requirements of time-critical services and should become a new direction for network resource optimization. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things, which can maximize system utility while ensuring deterministic latency.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] A resource allocation method for end-to-end deterministic latency network slicing assisted by digital twins in industrial Internet of Things, the method comprising the following steps:

[0008] S1: In the industrial IoT scenario, a digital twin DT-assisted end-to-end network slicing architecture is built. The digital twin is used to synchronize VNFs to perceive each network slice request information, including the source node, destination node, base station location information, virtual link set, node resource information, latency requirements, and reliability requirements of the service flow.

[0009] S2: Consider the impact of computing resource offset and synchronization delay in the digital twin mapping process on the end-to-end latency of the service;

[0010] S3: Stochastic network calculus (SNC) is introduced to characterize the upper bound of service transmission delay. The service arrival model and service model statistically analyzed by the digital twin model are used to characterize the relationship between the upper bound of delay and reliability.

[0011] S4: Jointly optimize the multi-dimensional resource allocation problem of time-frequency, computing, storage, and bandwidth with the goal of maximizing system utility;

[0012] S5: To solve the optimization problem, a digital twin-assisted multi-agent priority experience replay flexible actor-critic algorithm DT-PER-MASAC is proposed to learn the network slice deployment strategy that maximizes the system utility while ensuring the latency and reliability requirements of end-to-end network slices.

[0013] Furthermore, the DT-assisted end-to-end network slicing architecture includes a physical network layer, a digital twin layer, and an application service layer;

[0014] The physical network layer includes physical network elements, namely base stations (BS) in the Radio Access Network (RAN) domain, mobile edge computing (MEC) servers close to the BS, IoT terminals, and multiple servers in the Core Network (CN) domain. Physical network elements support both software and virtualization technologies to orchestrate and manage all network functions in the network.

[0015] The digital twin layer is composed of RAN and CN domain mappings, including multiple virtual devices, virtual nodes, and links. The digital twin layer assists the end-to-end network, predicts and analyzes transmission status based on real-time and historical data, and further adjusts resource allocation plans in the digital twin space.

[0016] The application service layer checks whether the service level agreement of the deployed end-to-end slices meets the requirements by interacting with the DT layer.

[0017] Furthermore, the network resources of the end-to-end network slice include heterogeneous communication resources, computing resources, storage resources and link bandwidth resources.

[0018] Furthermore, the service is provided by RAN slices. The RAN delay model includes uplink transmission delay and processing delay. The CN slice considers propagation delay, transmission delay, processing delay and DT synchronization delay. The propagation delay and transmission delay are calculated in the routing stage, and the processing delay and DT synchronization delay are calculated at the node.

[0019] Furthermore, in S3, the arrival process of service data packets within the network domain and the service process of service nodes are analyzed using SNC and moment generating functions to obtain the upper bound of the delay violation probability. Given the arrival traffic distribution and delay constraints, the delay and service transmission reliability are linked:

[0020] P[τ m,E2E <T m ]≥1-ε m,max

[0021] Where, τ m,E2E is the total end-to-end actual delay of service m in DTN, including RAN delay τ m,RAN and CN delay τ m,CN , denoted as τ m,E2E =τ m,RAN +τ m,CN , T m Indicates the end-to-end delay requirement of the guaranteed service m data packet, p m =1-εm,max Represents the service transmission reliability requirement, ε m,max Indicates the maximum transmission violation probability of service delay.

[0022] In the formula, the total end-to-end actual delay of service m in DTN is expressed as τ m,E2E =τ m,RAN +τ m,CN , T m Indicates the end-to-end delay requirement of the guaranteed service m data packet, p m =1-ε m,max Indicates the service transmission reliability requirement.

[0023] Furthermore, in S4, the problem of ensuring the determinism of the DT-assisted E2E network slice delay while maximizing the total system utility is expressed as:

[0024]

[0025] Where N RB Indicates the time-frequency resources allocated in the RAN domain, Indicates the computing resources on the server nodes in the RAN domain and CN domain, Indicates storage resources within the CN domain. represents the bandwidth resources between nodes, M represents the total number of services, Represents the total utility of the system.

[0026] Furthermore, in S5, the optimization problem is transformed into a Markov decision process model, including state, action, transition probability and reward, which is expressed as This includes state sets Action Set State transition probability set Bonus Set The agent initiates a service request for each IoT terminal in the system scenario, and each agent selects an action from the action space by considering the network status.

[0027] Furthermore, in S5, the DT-PER-MASAC algorithm includes the following steps:

[0028] S81: Initialize the actor network parameters of the agent, the evaluation network parameters and target network parameters in the critic network, and the experience replay pool;

[0029] S82: Determine whether the set number of iterations is exceeded, if so, stop the iteration, otherwise continue to execute S83;

[0030] S83: For N samples in the experience replay pool, obtain the initial local state information of each agent from the DT-assisted E2E network slicing environment;

[0031] S84: Determine whether the set number of training times is exceeded, if so, stop training, otherwise continue to execute S85;

[0032] S85: Update the actor network and critic network based on the priority experience replay mechanism;

[0033] S86: Recalculate the temporal difference-error (TD-error) value of the extracted experience, update the priority and temperature coefficient of the experience, and soft-update the target network parameters.

[0034] 9. The resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things according to claim 7 is characterized in that the Prioritized Experience Replay (PER) in the DT-PER-MASAC algorithm uses priority to calculate the probability of random selection, and the priority of sample i is expressed as:

[0035]

[0036] Among them, s t and s t+1 are the states of time slot t and the next time slot, a t and a t+1 are the actions selected in time slot t and the next time slot, r(s t ,a t ) is based on the state s t Execute action a t Rewards received, and are the action value functions of the current time slot and the next time slot respectively, α is the temperature coefficient, π φ (a t+1 |s t+1 ) indicates that the strategy is in state s t+1 Next take action a t+1 The probability of is a positive constant that approaches 0 infinitely, ensuring that every transition can be sampled even if the TD error is zero.

[0037] The beneficial effects of the present invention are: In response to the problem that cross-network domain resource allocation in end-to-end network slicing is difficult to meet the latency and reliability requirements due to the dynamic changes in business requests and service node status of IoT terminals, the present invention proposes a digital twin-assisted end-to-end deterministic latency network slicing resource allocation method. First, a DT-assisted end-to-end network slicing frame is constructed. Secondly, based on the random network calculus theory, the probability of end-to-end delay violations is analyzed, and a DT-assisted end-to-end resource allocation model is established to ensure deterministic delay while maximizing system utility. Finally, it is proposed to use a multi-agent deep reinforcement learning algorithm in a distributed architecture to solve complex optimization problems and achieve efficient network resource allocation. The present invention can improve system utility while meeting deterministic delay requirements.

[0038] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0040] Figure 1 Schematic diagram of the end-to-end network slicing architecture assisted by digital twins;

[0041] Figure 2 This is the network structure diagram of the DT PER-MASAC algorithm. DETAILED DESCRIPTION

[0042] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0043] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0044] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0045] An embodiment of the present invention proposes a resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things, which can maximize system utility while ensuring deterministic latency.

[0046] The steps of this method are as follows:

[0047] S1: In the industrial IoT scenario, a digital twin-assisted end-to-end network slicing architecture is built. The digital twin is used to synchronize VNFs to perceive each network slice request information, including the source node, destination node, base station location information, virtual link set, node resource information, latency requirements, and reliability requirements of the service flow. Figure 1 Schematic diagram of the end-to-end network slicing architecture assisted by digital twins.

[0048] The physical network layer includes physical network elements, such as base stations within the RAN domain, MEC servers located near the base stations, IoT terminal devices, and multiple servers within the CN domain. Each server contains one or more virtual machines, and servers are connected by optical fiber. Within the physical network layer, all physical network elements support both software-based and virtualized technologies to effectively orchestrate and manage all network functions, network resources, and network traffic data within the network.

[0049] The digital twin function is deployed on multiple CN servers, and determines the resource allocation strategy for the physical network by updating the wireless transmission channel time and frequency resources, server resource status, service request status, historical data, etc. in real time. The digital twin layer is composed of the RAN domain and CN domain mapping, including multiple virtual IoT terminal devices, an MEC server and multiple CN domain VNFs. More specifically, it includes the location information of the service terminal devices, service latency requirements and guaranteed transmission probability, link topology between VNF network elements, and bandwidth resources. The service m of the digital twin layer is represented as

[0050]

[0051] in, Represents the digital twin of IoT terminal m, including the terminal's location coordinate information and service SLA (latency, reliability). Represents the collection of MEC servers in the RAN domain and servers in the CN domain, including the server's computing and storage resource usage. It builds a service digital twin by extracting the operating status of multiple related devices.

[0052] In the DT-assisted end-to-end network slicing framework, the physical network layer and the digital twin layer continuously interact to implement data collection and slice control management. By mapping the characteristics and status of physical entities and communication environments, a more accurate real-time DT model can be constructed.

[0053] Building a fully mapped twin model of all physical network elements will consume a large amount of network resources. Based on the proposed digital twin network architecture, the digital twin synchronization and resource allocation process is as follows:

[0054] First, raw data is obtained from the physical network layer, such as device location, base station location, server available resources, service quality indicators, etc. After the raw data is collected, necessary preprocessing, such as data cleaning, is performed to ensure the effectiveness and accuracy of further analysis.

[0055] Secondly, the DT VNF collects synchronous data packets consisting of key information at a certain sampling rate and calculates the synchronization delay of transmitting synchronous data. It must ensure that the transmission of data packets remains within the maximum allowable synchronization delay to maintain the real-time performance of the DT model.

[0056] The DT VNF then stores and analyzes the received data, constructing or updating the DT model to reflect the real-time state of the physical network. Furthermore, within the constructed DT model, an AI algorithm model is used to formulate or adjust resource allocation plans for network slices based on the service's slicing requirements and current resource availability. This plan is then distributed to the virtual network for verification. The virtual network feeds the obtained synchronization delay back to the DT model to determine whether synchronization should be performed. If the synchronization delay observed in the virtual network exceeds the maximum synchronization delay, it may be due to significant deviations in the data or parameters collected in the DT model, requiring adjustments to the DT mapping model. This process is repeated and updated until the synchronization mapping requirements are met.

[0057] Finally, the resource allocation plan is sent to the physical network layer to guide configuration and adjustment within the physical network. The physical network layer and the digital twin layer continuously exchange information, forming a closed-loop control system, thereby achieving highly adaptive and optimized network management.

[0058] Computing capacity is increased in equipment rooms near the access network, enabling RAN virtualization through MEC servers. Virtual network functions are instantiated by allocating physical resources (CPU, RAM). In a single-cell, multi-service MEC system, a single base station (BS) is equipped with one high-performance MEC server and handles service requests from m IoT devices within range. The service set M = {1, 2, …, m}, where each service m∈M is transmitted using the OFDMA protocol. Service terminals transmit data packets to the BS via the uplink.

[0059] Each service is served by a RAN slice, and the total bandwidth of the RAN slice is B = J·B J , the frequency domain is divided into J sub-channels with a bandwidth of B J In the time domain, a scheduling frame can be divided into K scheduling subframes of length Δt. The minimum granularity of the time-frequency resource is the set of resource blocks (RBs) R = {1, 2, …, R}, and the total number of RBs is R = J·K.

[0060] A CN network slice is a virtual link connected by multiple VNF nodes. Different virtual links have different latency requirements, transmission probabilities, and resource requirements. CN is represented as an undirected graph. Represents the node set, E={1,2,…,e,…,E} represents the link set between nodes. Each server node n deploys one or more virtual machines, denoted as V n = {1, 2, ..., b}, where b represents the number of VMs hosted by node n. Assuming that each VM is dedicated to a VNF and VNFs cannot share the same VM resources, the VNF set on node n is represented as The total number of virtual machines in CN is

[0061] Different services are run by different slices. The network service provider allocates appropriate network resources to ensure deterministic latency, including computing resources, storage resources for VNFs on transmission links, and bandwidth resources for routing traffic between VNFs. The total resource capacity of each node n∈N is expressed as in The total capacity of computing resources and storage resources respectively. There is a limited bandwidth capacity between two nodes. i and j represent any two nodes in the node set N, where i is not equal to j. Different virtual links do not share the same VNF; each VNF uses exclusive network resources.

[0062] S2: Consider the impact of computing resource offset and synchronization delay in the digital twin mapping process on the end-to-end latency of the service;

[0063] Whether the RB resource r on the BS is allocated to service m is represented by a binary variable Indicates that a resource block can only be used for one service transmission:

[0064]

[0065] The number of RBs required for the transmission of service m is The amount of time and frequency resources used by all services cannot exceed the total amount of time and frequency resources:

[0066]

[0067] Although the DT model represents the operating state of the real network as accurately as possible, there are still mapping errors due to the limitations of the DT modeling method and modeling data acquisition. Considering that the digital twin mapping process has a computing resource offset Δf m , its range can be obtained by summarizing the long-term operation of DT, |Δf m |≤Δf max , Δf max is the maximum computing resource offset. When the computing resource offset exceeds the threshold Δf max When , it indicates that the current twin network element computing resource mapping error is large, and the computing resource offset needs to be corrected, that is, the DT mapping model needs to be synchronized. The computing resources actually allocated to service m are ψ CPU,m =ψ CPU,m +Δf m ,ψ CPU,m ≥0, the allocated computing resources cannot exceed the total computing resources of the MEC server:

[0068]

[0069] The RAN latency model includes uplink transmission latency and processing latency. The BS implements RAN slicing by allocating RBs. The total RAN slice latency for service m is expressed as τ m,RAN =t up,m +t comp,m .

[0070] The CN slice model of the transmission service m in the digital twin layer is represented as a virtual undirected graph in are the virtual nodes and virtual links of business m respectively. Contains parameters They represent the source node, destination node, delay requirement, and reliability requirement of the virtual link transmitting service m.

[0071] The physical network model is mapped to the digital twin layer, setting binary variables Indicates that the physical network node n is mapped to VNFu, a binary variable Indicates that the physical link (i, j) is mapped to the virtual link (u, v). Each physical node can only be mapped to one VNF node in the same virtual link, which is represented by

[0072] The computing and storage resources of each node are shared by multiple VNFs deployed on the node. Each VNF instance has exclusive resources and can only belong to a certain service link. The computing and storage resource constraints of node n are:

[0073]

[0074] in, Indicates the computing resources allocated to VNF u for service m, expressed in CPU cycles. Indicates the storage resources allocated to VNF u.

[0075] The bandwidth resources allocated between VNF u and v on the virtual link in the transport service m are The bandwidth resources allocated to virtual links should be smaller than the actual bandwidth resources between physical links. Link bandwidth resource constraints:

[0076]

[0077] A virtual link is dedicated to transmitting a single service type. Each slice has exclusive access to the resources it receives, preventing link congestion. Therefore, there is no queuing for service transmission, and queuing delay is not considered. The CN domain in the digital twin layer considers propagation delay, transmission delay, processing delay, and synchronization delay required for DTN synchronization. Propagation delay and transmission delay are calculated during the routing phase, while processing delay and synchronization delay are calculated at the node. The actual delay within the CN domain is expressed as:

[0078] τ m,CN =t prop,m +t trans,m +t proc,m +t syn,m

[0079] The node processing latency can be viewed as a function of the computing and storage resources allocated to VNF u on the physical node n. Considering the offset of computing resources between the twin network element and the real network element, the processing latency of VNF u on node n is expressed as:

[0080]

[0081] The DT model requires continuous synchronous data transmission with the modeled object to maintain consistency. The size of the synchronization data packet is relatively small in actual transmission. Resource management and scheduling require real-time information updates. If the node synchronization delay is large, resource scheduling will not be timely, thereby increasing the end-to-end delay. However, data packet loss may occur in the communication between the DT server and the physical network element nodes. In order to measure the state consistency between the DTN and its physical network elements, the DT synchronization delay is introduced. The DT synchronization delay is the maximum synchronous data transmission delay between the DT model and the mapped server node set:

[0082]

[0083] in, Indicates the time required to send one unit of data per unit distance, D syn,m Indicates the size of the synchronization data packet, d i,u Represents the distance between DT synchronization node i and VNF u. The maximum synchronization delay defines the maximum allowable time for data to be transmitted from the physical network layer to the digital twin layer to complete synchronization. By setting the maximum synchronization delay, it can be ensured that the DT model can reflect the current status of the physical network in a timely manner. The DT synchronization delay should be less than the maximum synchronization delay, which is expressed as t syn,m ≤t syn,max .

[0084] The total end-to-end actual delay of service m in DTN is τ m,E2E =τ m,RAN +τ m,CN , the end-to-end delay constraint of service m is:

[0085] τ m,E2E ≤T m

[0086] S3: Stochastic network calculus is introduced to characterize the upper bound of service transmission delay. The service arrival model and service model statistically analyzed by the digital twin model are used to characterize the relationship between the upper bound of delay and reliability.

[0087] SNC is a theoretical tool for network performance analysis that can guarantee the network's delay boundary performance within a certain probability range. It is very suitable for business transmission with strict delay requirements while taking into account resource utilization. SNC uses the arrival curve and service curve to constrain the arrival process of business data and the service process of service nodes, and can further analyze the backlog limit and delay limit. The backlog limit is the maximum vertical distance between the two curves, and the delay limit is the maximum horizontal distance between the two curves. For business m, all data packets have the same delay requirement T m , ε m,maxis the maximum delay violation probability of service m. If the service data packet has a first-in-first-out (FIFO) policy, it is guaranteed that the service m data packet meets the end-to-end delay requirement T m , reliability requirement p m =1-ε m,max The lower transmission is expressed as

[0088] P[τ m,E2E <T m ]≥1-ε m,max

[0089] In the service flow arrival model, service flows maintain independent and identical distribution. The process of service data packets from IoT device terminals to the BS is a Poisson process. The arrival process of service m arriving within the time [τ, t] is:

[0090]

[0091] Among them, a m (i) represents the data packet size of service m from the service terminal to the BS in the i-th time slot. Given a constant data packet size D m , the cumulative arrival process is A(t)=D m N(t), calculate A m The moment generating function of (τ,t) is

[0092]

[0093] Among them, the free parameter θ>0, for the SNC problem, the affine reaches the envelope upper bound

[0094]

[0095] Using affine to reach envelope parameters and Restricted Arrival Process A m (τ,t):

[0096]

[0097] The wireless channel is modeled as a Rayleigh fading channel, and the instantaneous signal-to-noise ratio of service m transmitted in time slot i is recorded as The probability density function of g follows an exponential distribution with a mean of 1. If the channel gain of each time slot is independent and identically distributed, then the service process is also independent and identically distributed, and the channel transmission service process is The moment generating function is:

[0098]

[0099] make Get the channel transmission service process moment generating function:

[0100]

[0101] Get the affine service envelope parameters:

[0102]

[0103] Service data is transmitted to CN via BS, and processed by multiple server nodes. Server node i is modeled as having a total computing capacity of C i ,φ m ∈[0,1] represents the resource allocation weight factor for business m. The service calculation process of server node i is: The moment generating function of the server computing service is get

[0104] The service process S(t) of the business data packet is represented by two parts: the wireless transmission service process and the server computing service process. According to the service cascade theorem, The service node i∈(1,n) is calculated, and n represents the number of service nodes on the transmission link of service m. Let the dynamic server S0 represent the service of the wireless channel, and the dynamic servers S1,…,S n Represents the service of the server on the transmission link, free parameter ∈>0 and

[0105] When each service process is independent of each other, the end-to-end network service process constraint can be obtained as follows:

[0106]

[0107] Under the condition that the arrival process and the service process are independent, the upper bound of the delay violation probability is obtained:

[0108]

[0109] Different service transmissions have different reliability requirements. The reliability requirement of service m is expressed as the relationship between the upper bound of the end-to-end delay violation probability and the maximum transmission violation probability of the service delay:

[0110] ε m <ε m,max

[0111] S4: Jointly optimize the multi-dimensional resource allocation problem of time-frequency, computing, storage, and bandwidth with the goal of maximizing system utility;

[0112] NSP accepts the request for slice m and allocates data rate R m Generate income, represents the data rate unit price of slice m, and the revenue function is expressed as:

[0113]

[0114] use They represent the unit cost of allocating time-frequency resource blocks in the RAN domain and the unit cost of computing resources in the MEC server for slice m, respectively. The cost function of slice m in the RAN domain is:

[0115]

[0116] represents the unit cost of computing resources of nodes on slice m in the CN domain, Indicates the unit cost of storage resources of the node, It represents the unit cost of bandwidth resources on the transmission link. The cost function of slice m in the CN domain is:

[0117]

[0118] The total cost function of DT-assisted E2E network slicing is expressed as The overall utility function is expressed as:

[0119]

[0120] Among them, Θ1 and Θ2 are used to balance and scale the revenue and cost of different service types respectively. In order to ensure the determinism of the DT-assisted E2E network slicing latency while maximizing the total utility of the entire system, the optimization objective is expressed as:

[0121]

[0122] S5: To solve the optimization problem, a digital twin-assisted multi-agent priority experience replay flexible actor-critic algorithm DT-PER-MASAC is proposed to learn the network slice deployment strategy that maximizes the system utility while ensuring the latency and reliability requirements of end-to-end network slices. Figure 2 This is the network structure diagram of the DT PER-MASAC algorithm.

[0123] The optimization problem is transformed into a Markov decision process model, including states, actions, transition probabilities, and rewards, which can be expressed as This includes state sets Action Set State transition probability set Bonus Set The agent initiates a service request for each IoT terminal in the system scenario, and each agent selects an action from the action space by considering the network status;

[0124] The Soft Actor-Critic (SAC) algorithm is an offline policy algorithm based on a continuous action space and designed for the maximum entropy RL framework. Its optimal strategy is to maximize its entropy-regularized reward, making the algorithm more stable. The SAC algorithm objective function not only learns a policy that maximizes the expected cumulative reward, but also requires that the entropy of each action output by the policy be maximized:

[0125]

[0126] Where T is the number of time steps that the agent interacts with the environment in each round, ρπ represents the strategy π(a t |s t ), α is a hyperparameter of the temperature coefficient, which is used to determine the importance of entropy.

[0127] Update the actor network according to minimizing the KL divergence, in order to minimize the optimization target J using the stochastic gradient method. π (Θ), the sampling process is moved out of the computational graph by re-parameterization, and the state and random noise are used as inputs. t =f θ (ε t ,s t ), noise ε t Obey the standard normal distribution. We get:

[0128]

[0129] To automatically update α, represents the hyperparameter of the target entropy, and the optimized loss function is:

[0130]

[0131] SAC uses two Q networks and selects the smaller Q value as the target Q value each time to ensure fast and stable training, and updates the critic network by minimizing the Bellman error.

[0132] The SAC algorithm is only applicable to a single agent. If the entire system is modeled and then centralized, single-agent reinforcement learning is employed, the action space becomes large and signaling overhead is high. Furthermore, if independent reinforcement learning is employed for each agent, the strategies of other agents may change at any time while a particular agent is being trained, leading to an unstable environment. This approach employs a multi-agent approach to apply the SAC algorithm to a multi-agent framework with centralized training and distributed execution, enabling collaborative efforts between agents to solve optimization problems.

[0133] The experience replay buffer typically uses uniform sampling to sample data. This random update results in the use of a large amount of ineffective or invalid data. Furthermore, since each update uses uniform sampling, the algorithm converges slowly, and some rare but important data is not captured during the entire update process.

[0134] Using the PER structure, all data is sorted by priority (time difference, TD error). Data with higher priority, i.e., data with larger TD error, is more likely to be selected. Therefore, compared to a typical replay buffer, PER has a greater chance of selecting valid data in each time slot of the algorithm update. This results in a larger algorithm update gradient and faster algorithm convergence. PER uses priority to calculate the probability of random selection. The priority of sample i is expressed as:

[0135]

[0136] Among them, s t and s t+1 are the states of time slot t and the next time slot, a t and a t+1 are the actions selected in time slot t and the next time slot, r(s t ,a t ) is based on the state s t Execute action a t Rewards received, and are the action value functions of the current time slot and the next time slot respectively, α is the temperature coefficient, π φ (a t+1 |s t+1 ) indicates that the strategy is in state s t+1 Next take action a t+1 The probability of is a positive constant that approaches 0 infinitely, ensuring that every transition can be sampled even if the TD error is zero.

[0137] The probability of sampling i is defined as κ is the sampling preference control parameter, which is used to determine the sampling preference between uniform sampling and greedy sampling. When κ = 0, it is uniform sampling, and when κ = 1, it is greedy sampling.

[0138] The priority distribution in the experience replay pool is constantly changing, and the correlation between samples may change the results of algorithm convergence.

[0139] In order to ensure that there is no correlation between samples, the importance sampling method is used. The importance sampling weight of sample i is defined as:

[0140]

[0141] Where ζ is the importance weight control parameter, which is used to determine the extent to which the impact of the priority replay mechanism on the convergence result is offset.

[0142] Calculate the experience extraction probability P(i) and importance sampling weight w i After that, the loss function of the updated Critic network becomes:

[0143]

[0144] The DT-PER-MASAC algorithm specifically includes the following steps:

[0145] S81: Initialize the actor network parameters of the agent, the evaluation network parameters and target network parameters in the critic network, and the experience replay pool;

[0146] S82: Determine whether the set number of iterations is exceeded, if so, stop the iteration, otherwise continue to execute S83;

[0147] S83: For N samples in the experience replay pool, obtain the initial local state information of each agent from the DT-assisted E2E network slicing environment;

[0148] S84: Determine whether the set number of training times is exceeded, if so, stop training, otherwise continue to execute S85;

[0149] S85: Update the actor network and critic network based on the priority experience replay mechanism;

[0150] S86: Recalculate the TD-error value of the extracted experience, update the priority and temperature coefficient of the experience, and soft-update the target network parameters.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A resource allocation method for end-to-end deterministic latency network slicing in the industrial Internet of Things (IIoT) assisted by digital twins, characterized by: The method comprises the following steps: S1: In the industrial IoT scenario, a digital twin DT-assisted end-to-end network slicing architecture is built. The digital twin is used to synchronize VNFs to perceive each network slice request information, including the source node, destination node, base station location information, virtual link set, node resource information, latency requirements, and reliability requirements of the service flow. S2: Consider the impact of computing resource offset and synchronization delay in the digital twin mapping process on the end-to-end latency of the service; S3: Stochastic network calculus (SNC) is introduced to characterize the upper bound of service transmission delay. The service arrival model and service model statistically analyzed by the digital twin model are used to characterize the relationship between the upper bound of delay and reliability. S4: Jointly optimize the multi-dimensional resource allocation problem of time-frequency, computing, storage, and bandwidth with the goal of maximizing system utility; S5: To solve the optimization problem, a digital twin-assisted multi-agent priority experience replay flexible actor-critic algorithm DT-PER-MASAC is proposed to learn the network slice deployment strategy that maximizes the system utility while ensuring the latency and reliability requirements of end-to-end network slices.

2. The resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things according to claim 1 is characterized by: The DT-assisted end-to-end network slicing architecture includes the physical network layer, the digital twin layer, and the application service layer; The physical network layer includes physical network elements, namely base stations (BSs) in the radio access network (RAN) domain, mobile edge computing (MEC) servers close to the BSs, IoT terminals, and multiple servers in the core network (CN) domain. Physical network elements support both software and virtualization technologies to orchestrate and manage all network functions in the network. The digital twin layer is composed of RAN and CN domain mappings, including multiple virtual devices, virtual nodes, and links. The digital twin layer assists the end-to-end network, predicts and analyzes transmission status based on real-time and historical data, and further adjusts resource allocation plans in the digital twin space. The application service layer checks whether the service level agreement of the deployed end-to-end slices meets the requirements by interacting with the DT layer.

3. The resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things according to claim 2 is characterized by: The network resources of the end-to-end network slice include heterogeneous communication resources, computing resources, storage resources and link bandwidth resources.

4. The resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things according to claim 1 is characterized by: The service is provided by RAN slices. The RAN latency model includes uplink transmission delay and processing delay. The CN slice considers propagation delay, transmission delay, processing delay and DT synchronization delay. Propagation delay and transmission delay are calculated in the routing stage, and processing delay and DT synchronization delay are calculated at the node.

5. The method for allocating resources for digital twin-assisted end-to-end deterministic latency network slices in the industrial Internet of Things according to claim 1 is characterized by: In S3, the SNC and moment generating function are used to analyze the arrival process of service data packets in the network domain and the service process of service nodes, obtain the upper bound of the delay violation probability, and establish a connection between delay and service transmission reliability under given arrival traffic distribution and delay constraints: P[t m,E2E <T m ]≥1-e m,max Where, τ m,E2E is the total end-to-end actual delay of service m in DTN, including RAN delay τ m,RAN and CN delay τ m,CN , denoted as τ m,E2E =τ m,RAN +τ m,CN , T m Indicates the end-to-end delay requirement of the guaranteed service m data packet, p m =1-ε m,max Represents the service transmission reliability requirement, ε m,max Indicates the maximum transmission violation probability of service delay.

6. The resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things according to claim 1 is characterized by: In S4, the problem of ensuring the determinism of the DT-assisted E2E network slice latency while maximizing the total system utility is expressed as: Where N RB Indicates the time-frequency resources allocated in the RAN domain, Indicates the computing resources on the server nodes in the RAN domain and CN domain, Indicates storage resources within the CN domain. represents the bandwidth resources between nodes, M represents the total number of services, Represents the total utility of the system.

7. The resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things according to claim 1 is characterized by: In S5, the optimization problem is transformed into a Markov decision process model, including states, actions, transition probabilities and rewards, which can be expressed as This includes state sets Action Set State transition probability set Bonus Set The agent initiates a service request for each IoT terminal in the system scenario, and each agent selects an action from the action space by considering the network status.

8. The resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things according to claim 1 is characterized by: In S5, the DT-PER-MASAC algorithm includes the following steps: S81: Initialize the actor network parameters of the agent, the evaluation network parameters and target network parameters in the critic network, and the experience replay pool; S82: Determine whether the set number of iterations is exceeded, if so, stop the iteration, otherwise continue to execute S83; S83: For N samples in the experience replay pool, obtain the initial local state information of each agent from the DT-assisted E2E network slicing environment; S84: Determine whether the set number of training times is exceeded, if so, stop training, otherwise continue to execute S85; S85: Update the actor network and critic network based on the priority experience replay mechanism; S86: Recalculate the time difference error TD-error value of the extracted experience, update the priority and temperature coefficient of the experience, and soft-update the target network parameters.

9. The resource allocation method for digital twin-assisted end-to-end deterministic latency network slicing in the industrial Internet of Things according to claim 7 is characterized by: The priority experience replay PER in the DT-PER-MASAC algorithm uses priority to calculate the random selection probability of the sample. The priority of sample i is expressed as: Among them, s t and s t+1 are the states of time slot t and the next time slot, a t and a t+1 are the actions selected in time slot t and the next time slot, r(s t ,a t ) is based on the state s t Execute action a t Rewards received, and are the action value functions of the current time slot and the next time slot respectively, α is the temperature coefficient, π φ (a t+1 |s t+1 ) indicates that the strategy is in state s t+1 Next take action a t+1 The probability of is a positive constant that approaches 0 infinitely, ensuring that every transition can be sampled even if the TD error is zero.