Digital twin-driven Internet of Things user behavior prediction and resource allocation method

By building a digital twin-driven edge IoT network slicing system, combined with SA-TCN and PER-MATD3 algorithms, the resource allocation lag and load imbalance of traditional network architectures under high mobility terminals and time-varying business needs are solved, high-precision prediction and dynamic resource management are achieved, and system utility and resource utilization are improved.

CN120238922APending Publication Date: 2025-07-01CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510595895.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When traditional network architectures face high mobility terminals and time-varying business needs, resource allocation lag, load imbalance and end-to-end delays increase, resulting in a decline in service quality and low resource utilization. The existing digital twin methods lack generalization capabilities in dynamic network environments, making it difficult to achieve accurate prediction and dynamic management.

Method used

Build a digital twin-driven edge IoT network slicing system, combine the self-attention mechanism and the time convolutional network (SA-TCN) model to predict user behavior, and improve the generalization ability of the model through meta-learning, combine the PER-MATD3 algorithm to optimize resource allocation, and establish end-to-end delay, DT reliability and system cost models to maximize system effectiveness.

Benefits of technology

It significantly improves user behavior prediction accuracy and real-time resource scheduling, reduces end-to-end delay, balances network load, improves resource utilization and system comprehensive utility, and adapts to high-mobile terminals and time-varying business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238922A_ABST
    Figure CN120238922A_ABST
Patent Text Reader

Abstract

The invention relates to a digital twin-driven Internet of Things user behavior prediction and resource allocation method, and belongs to the field of mobile communication. Aiming at the problems of service quality reduction, load imbalance, low utilization rate and the like caused by resource allocation lag in the traditional method, the invention provides the following scheme: constructing a three-level system model of a physical infrastructure layer, a digital twin network layer and a control layer; an SA-TCN model combining a self-attention mechanism and a time convolution network is designed, and the prediction precision in a dynamic environment is improved through meta learning; establishing an end-to-end delay, DT credibility and system cost model, and taking maximization of system utility as a target; and dynamically optimizing a resource allocation strategy by adopting a PER-MATD3 algorithm. According to the method, high-precision user behavior prediction is realized, the real-time performance of resource scheduling is improved, the resource utilization rate, the load balance and the comprehensive effectiveness of the system are remarkably improved under the condition that end-to-end delay and credibility constraints are guaranteed, and the method is suitable for a high-mobility Internet of Things scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of mobile communications and relates to a method for predicting Internet of Things user behavior and allocating resources driven by digital twins. Background Art

[0002] The rapid development of Internet of Things technology has promoted the widespread application of smart devices. However, high-mobility terminals and time-varying service requirements pose severe challenges to the network architecture. Frequent changes in device locations, dynamic fluctuations in service requests, and real-time requirements for resource allocation result in significant deficiencies in the traditional network architecture in terms of resource utilization and service quality guarantee. Due to relying on passive resource allocation strategies, traditional methods are difficult to adapt to dynamic environments and often cause problems such as lagging resource allocation, load imbalance, and increased end-to-end delay, thereby affecting network stability and user experience.

[0003] As the core architecture of 5G, network slicing divides the physical network into multiple virtual subnets to provide customized services for different business scenarios. Combining edge computing technology, network slicing can enhance the flexibility of resource scheduling. However, in scenarios with time-varying mobile users and service requirements, accurately predicting user behavior and achieving dynamic resource management remain difficult. The behavior of Internet of Things users has strong temporal characteristics and high dynamics. Traditional scheduling algorithms are limited by real-time data processing capabilities and model generalization capabilities and cannot effectively capture complex temporal dependencies, resulting in insufficient prediction accuracy and lagging resource allocation strategies.

[0004] Digital twin technology provides a new idea for network optimization by constructing a virtual mapping of physical devices and synchronizing state data in real time. By integrating device historical behavior, real-time status, and environmental information, it can enhance the feature expression ability of temporal prediction, thereby supporting proactive resource allocation. However, existing digital twin-based methods face problems such as insufficient model generalization ability and poor multi-task adaptability in dynamic network environments. In addition, the high mobility of Internet of Things terminals exacerbates the state deviation between the digital twin model and physical devices, and synchronization delay and resource migration cost further restrict the system utility.

[0005] Therefore, there is an urgent need for an innovative method that integrates efficient temporal prediction and dynamic resource scheduling to solve the deficiencies of traditional architectures in prediction accuracy, real-time response, and resource utilization, so as to meet the requirements of the Internet of Things for low latency, high reliability, and efficient resource utilization. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a digital twin-driven Internet of Things user behavior prediction and resource allocation method, which solves problems such as the decline in service quality, load imbalance, and low resource utilization rate caused by lagging resource allocation strategies or uneven resource allocation, and maximizes the system utility under the constraints of end-to-end delay and DT credibility. At the same time, by predicting the user behavior information of IoT terminals, it analyzes and formulates resource allocation strategies in advance to ensure the timeliness of resource allocation strategies and the stability of network performance.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A digital twin-driven Internet of Things user behavior prediction and resource allocation method, which includes the following steps:

[0009] S1. Construct a digital twin-driven edge Internet of Things network slice system model, which includes a physical infrastructure layer, a digital twin network layer, and a control layer;

[0010] S2. Establish an SA-TCN time series prediction model to predict user behavior information. This model enhances the ability to capture the dependence relationship of time series data features between different time steps through the self-attention mechanism, so as to effectively predict user behavior information (location coordinates and service resource requirements); during the training process, combined with the model-agnostic meta-learning method, the prediction model has stronger generalization ability in a time-varying network environment, and improves its training efficiency and prediction accuracy on different tasks;

[0011] S3. Establish an end-to-end delay model, a DT credibility model, a penalty / incentive model, and a system cost model, as well as an optimization goal of maximizing the system utility according to the digital twin-driven edge Internet of Things network slice system model;

[0012] S4. According to the prediction results of the user behavior information of IoT terminals, transform the resource allocation problem into a Markov decision process, and use the multi-agent double-delay deep deterministic policy gradient algorithm combined with prioritized experience replay to solve the resource allocation strategy with the optimization goal of maximizing the system utility.

[0013] Further, in S1, the physical infrastructure layer is mainly composed of an access network and a core network. The access network includes a base station, the MEC server of the base station, and IoT terminals. The sets of IoT terminals and base stations are respectively represented as and Each base station includes an MEC server; the core network is formed by connecting multiple servers to form a fully connected undirected graph G P ={N P , L P}.

[0014] The digital twin network layer realizes the twin mapping of the physical network by collecting the status information of the physical infrastructure. The DT network is an undirected graph G DT ={N DT , L DT}. The DT network analyzes the real-time and historical network status data of IoT terminals to predict their future behaviors (such as location movement and service requests). Based on the prediction results, it dynamically optimizes the resource allocation scheme. When the new scheme is better than the current configuration, it is sent and executed in real time through the SDN controller. This proactive predictive resource scheduling mechanism effectively avoids the lag of traditional passive allocation and significantly reduces the network trial-and-error cost and operation and maintenance risks.

[0015] Network slicing provides service support for the services of IoT terminals. Services can be abstracted as SFCs, and an SFC is a series of VNFs executed in sequence to meet specific service requirements. A set of SFCs represents different types of services, is the set of VNFs on at time t the demand for computing, storage, and bandwidth resources; The resource demand at time t is the sum of the resources required for all IoT terminal service requests it serves. Therefore, the service request of IoT terminals at time t can be represented as a set of VNFs Introduce a binary variable to represent and whether they are associated.

[0016] Furthermore, in S2, a self-attention TCN prediction model combined with the meta-learning method is used to predict future user behavior information. The user behavior information includes the location coordinate information of IoT terminals and the resource demand information of service requests. SA-TCN introduces a self-attention mechanism to characterize the correlation between different time steps and uses the self-attention mechanism to assign different weights to different elements of time series data.

[0017] First, for the input time series data adopt the sine position encoding P e , and input the encoded time series data X I ′ into the self-attention layer to calculate the attention scores at each time step, weight different time steps, calculate the attention scores through the query matrix Q, key matrix K, and value matrix V, and normalize the attention scores using the softmax function. The calculation formula is as follows:

[0018]

[0019] The time series data X″ after being processed by the self-attention layer I As the input of the TCN layer, the receptive field of the model is improved through the dilated causal convolution in multiple layers of TCN. The output of each layer of convolution is added to the input through a residual connection, and the ReLU activation function is applied to increase non-linearity. Finally, the pooling layer and the fully connected layer are used to map the feature information to the prediction space to obtain the final prediction result

[0020] The training process of the model-agnostic meta-learning method is divided into two stages: inner-layer training and outer-layer optimization. Through meta-training with alternating inner and outer layers on multiple different tasks, the SA-TCN prediction model can learn an optimized initial parameter. Thus, when facing a new network environment, it can quickly adapt with only a small number of samples and few gradient updates, improving the generalization ability of the prediction model. In the meta-training stage, the dataset of task T i is divided into a training set and a test set The training set is used for inner-layer training to optimize the prediction model parameters for a specific task; the test set is used for outer-layer optimization. After the inner-layer training is completed, it is used to evaluate the generalization ability of the prediction model and calculate the total test loss of all tasks to optimize the initial parameters of the model

[0021] The loss function of the inner-layer training can be expressed as:

[0022]

[0023] where O is the total number of samples in the training set; and are the input features and the corresponding true prediction target of the o-th sample in the training set of task T i respectively; is the prediction result of the prediction model based on the initial parameter θ on the o-th training set sample

[0024] The gradient update process of the inner-layer training can be expressed as:

[0025]

[0026] where α intra is the learning rate of the inner-layer training of the task; is the gradient of the loss function

[0027] In the outer-layer optimization, for task T i , its test loss function can be expressed as:

[0028]

[0029] where Z is the total number of samples in the test set​ and are respectively the input feature and the corresponding true prediction target of the z-th sample in the test set of task T i ; is the prediction result of the prediction model based on the specific task parameter θ i ' on the z-th test set sample.

[0030] The total test loss of all tasks can be expressed as:

[0031]

[0032] where I' is the total number of tasks.

[0033] Finally, the initial parameter θ of the prediction model is updated through outer optimization, and the update process can be expressed as:

[0034]

[0035] where β extra is the learning rate of outer optimization; is the gradient of the total test loss with respect to the initial parameter θ.

[0036] Furthermore, in S3, an end-to-end delay model, a DT credibility model, a penalty / incentive model, and a system cost model are constructed, and an optimization objective for maximizing system utility is constructed.

[0037] In the end-to-end delay model, the delay of a service request can be divided into wireless transmission delay and wired transmission delay. At time t, the wireless transmission delay of the uplink of IoT terminal u can be expressed as:

[0038]

[0039] where SINR u (t) is the signal-to-interference-plus-noise ratio of the transmitted signal; D u (t) is the size of the transmission data of the service request of IoT terminal u; B u (t) is the uplink bandwidth allocated by the base station to IoT terminal u.

[0040] The wired transmission delay is divided into the transmission delay between the MEC and the core network and the transmission delay in the core network. Since the delay of accessing the core network is much higher than the local delay, the transmission delay between the MEC and the core network is set to a fixed value T MC , and the transmission delay in the core network is divided into the processing delay of physical nodes and the communication delay between physical links. The processing delay can be expressed as:

[0041]

[0042] Among them, represents on the processing rate; represents the processing rate coefficient of service request data.

[0043] At time t, the communication delay can be expressed as:

[0044]

[0045] Among them, the binary variables and respectively represent the mapping relationships between the VNF node and the server, and between the virtual link and the physical link; represents the physical distance of the physical link l nn′ ; c represents the speed of light.

[0046] Due to the dynamic changes of service requests, the SFC needs to dynamically schedule VNFs to ensure service quality, that is, for VNF migration. The VNF migration process will generate a certain migration delay, thus affecting the end-to-end delay of the service. At time t, the total migration delay of can be expressed as;

[0047]

[0048] Among them, is the VNF migration data size; is the delay required to transmit a unit of data along the migration path; Ω is a positive coefficient; is the network routing hop distance between server nodes. Then the end-to-end delay of the service can be expressed as:

[0049]

[0050] In the DT credibility model, there are mainly two factors affecting the DT credibility of IoT terminals, namely DT mapping deviation and DT synchronization delay deviation. The DT mapping deviation can be expressed as:

[0051] DT map = ω c Δf c (t) + ω m Δf m (t)

[0052] Among them, ω c and ω m are the weight parameters of the CPU frequency deviation and the storage capacity deviation; Δf c (t) is the CPU frequency deviation, and Δf m (t) is the storage capacity deviation.

[0053] The DT deployment of the IoT terminal u is on the associated MEC server M u (t). The DT synchronization delay deviation mainly considers the impact of wireless transmission delay and MEC server processing delay. The MEC server processing delay is related to the computing resources of the server and can be expressed as: Related, it can be expressed as:

[0054]

[0055] Among them, Is the DT synchronization data size; Is the packet processing coefficient of the MEC server. Then the DT synchronization delay deviation is

[0056] Then for the IoT terminal u, its DT credibility can be expressed as:

[0057]

[0058] Among them, ψ DT ∈(0,1], the closer its value is to 1, the higher the DT credibility; ω map And ω syn Are weight parameters that determine the influence degree of mapping deviation and synchronization delay deviation on credibility.

[0059] In the penalty / incentive model, for the physical node n, its penalty function at time t can be expressed as:

[0060]

[0061] Among them, η n Is the node resource utilization rate; Are the lower and upper limits of the threshold of the resource utilization rate of the physical node; Is a fixed incentive for the node with balanced resource utilization rate.

[0062] In the system cost model, consider the DT migration cost and the VNF migration cost. DT migration occurs when M u (t - 1) ≠ M u (t). Then at time t, the total DT migration cost of the system can be expressed as:

[0063]

[0064] Among them, Represents the single DT migration cost, κ DT Is the DT migration cost factor; Is the migration data size of the DT; is the physical distance between MEC servers m and m'.

[0065] At time t, the total cost of VNF migration in the system can be expressed as:

[0066]

[0067] where is the cost of a single VNF migration, and κ VNF is the VNF migration cost factor. So at time t, the total cost of the system is C total = C DTmig (t) + C VNFmig (t)

[0068] The optimization goal of the system is to maximize the system utility, which can be expressed as:

[0069]

[0070] where C1 - C2 means that a VNF or a virtual link can only be mapped to a physical node or a physical link; C3 means that an IoT terminal will only be associated with the MEC server with the shortest distance; C4 is the end-to-end delay constraint, and is the maximum end-to-end delay constraint; C5 represents the DT credibility constraint of the IoT terminal, and is the maximum credibility constraint; C6 - C8 mean that the computing, storage, and bandwidth resource requirements of the VNF must be within the physical resource limits; C9 - C11 are binary variables.

[0071] Furthermore, in S4, based on the prediction result of the IoT terminal user behavior information, a network resource allocation strategy with the optimization goal of maximizing the system utility is formulated. Since traditional optimization methods may not be able to effectively handle each constraint condition in the optimization goal, the optimization goal is transformed into a Markov policy process, and the PER-MATD3 algorithm is used to solve the resource allocation strategy.

[0072] Regarding a network slice as an agent, assuming there are K agents, the agents maximize the system utility by optimizing the allocation of resources such as base stations and servers. The MDP model is established as including the state space the action space A and the reward function R;

[0073] The global reward of the agent can be expressed as:

[0074]

[0075] The basic framework of the PER-MATD3 algorithm is the Actor-Critic structure. Each agent k includes an Actor network with parameter λ k and two Actor networks with parameters μ k,1 and μk,2 The Critic network and its corresponding target network, with the target network parameters being λ′ k , μ′ k,1 , μ′ k,2 .

[0076] The prioritized experience replay mechanism can prioritize according to the importance of each sample, sampling high-priority samples more frequently from the experience pool, thereby improving the learning efficiency and accelerating the convergence speed. For agent k, for the i-th sample {S t , A t , R t , S t+1} in the experience replay pool, the priority Λ i is related to the TD error and can be expressed as:

[0077]

[0078] where γ represents the discount factor; is the Q value calculated by the target Critic network, and A t+1 is determined by the target Actor network; is a very small positive constant to ensure that even when the TD error is zero, the sample can still be sampled.

[0079] Since the PER mechanism disrupts the uniform sampling distribution of the original experience pool and may introduce bias, importance sampling weights Θ i are used for correction and can be expressed as:

[0080]

[0081] where, is a hyperparameter; is the total number of samples in the experience replay pool; P(i) is the sampling probability of the i-th sample.

[0082] The goal of the Critic network is to minimize the TD error, and the target loss function of the Critic network can be expressed as:

[0083]

[0084] where, represents the TD target value, and there is

[0085] Then, the gradient descent method is used to update the parameters of the Critic network:

[0086]

[0087] where β criticis the learning rate of the Critic network.

[0088] The goal of the Actor network is to learn the optimal policy. The Actor network approximates and optimizes the policy by maximizing the Q-value evaluated by the Critic network, so that the actions it selects can obtain the maximum Q-value. The objective loss function of the Actor network can be expressed as:

[0089]

[0090] Then, the gradient ascent method is used to update the parameters of the Actor network:

[0091]

[0092] where β actor is the learning rate of the Actor network.

[0093] The parameter updates of the target Actor network and the target Critic network can be expressed as:

[0094] λ′ k ← τλ k +(1 - τ)λ′ k

[0095] μ′ k,j ← τμ k,j +(1 - τ)μ′ k,j j = 1, 2

[0096] where τ is the update rate of the target network.

[0097] The beneficial effects of the present invention are as follows:

[0098] (1) By combining the self-attention mechanism with the temporal convolutional network (SA-TCN), the global dependence and local dynamic capture ability of the temporal characteristics of user behavior are enhanced, the prediction accuracy of the position coordinates and service resource requirements is significantly improved, and a reliable basis is provided for resource pre-allocation.

[0099] (2) By introducing the model-agnostic meta-learning method, the prediction model can quickly adapt with only a small number of samples in a time-varying network environment, solving the problem of insufficient generalization ability of traditional models and ensuring the stability of the prediction results in mobile scenarios.

[0100] (3) Based on the active prediction mechanism of the digital twin network, combined with the PER-MATD3 algorithm to achieve multi-agent collaborative decision-making, dynamically optimize the allocation of computing, storage, and bandwidth resources, effectively avoid the lag of traditional passive strategies, and reduce the end-to-end delay.

[0101] (4) Through the end-to-end delay model, DT credibility constraint, and cost-punishment incentive mechanism, while ensuring service quality, the load is balanced, the risk of resource overload or underload is reduced, and the migration cost of VNF and digital twin is lowered, maximizing the comprehensive utility of the system.

[0102] (5) The digital twin technology synchronizes the physical device status and virtual model in real time, and combines the prediction results to dynamically adjust the network slicing strategy, significantly improving the network's response ability to high-mobility terminals and time-varying service requirements, and ensuring service continuity.

[0103] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings

[0104] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0105] Figure 1 is the architecture diagram of the digital twin-driven edge Internet of Things network slice;

[0106] Figure 2 is the structure diagram of the SA-TCN prediction model;

[0107] Figure 3 is the network structure diagram of the PER-MATD3 algorithm. Detailed Embodiment

[0108] The following illustrates the implementation manners of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0109] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; for better illustrating the embodiments of the present invention, some components in the drawings will be omitted, enlarged, or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0110] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0111] Please refer to Figures 1 to 3 , a method for predicting IoT user behavior and resource allocation driven by digital twins.

[0112] A method for predicting IoT user behavior and resource allocation driven by digital twins proposed in this embodiment includes the following steps:

[0113] S1. Build a digital twin-driven edge IoT network slice system model, which includes a physical infrastructure layer, a digital twin network layer, and a control layer;

[0114] S2. Establish an SA-TCN time series prediction model to predict user behavior information. This model enhances the ability to capture the dependence relationship of time series data features between different time steps through the self-attention mechanism, so as to effectively predict user behavior information (location coordinates and service resource requirements); during the training process, combined with the model-agnostic meta-learning method, the prediction model has stronger generalization ability in a time-varying network environment, and improves its training efficiency and prediction accuracy on different tasks;

[0115] S3. Establish an end-to-end delay model, a DT credibility model, a penalty / incentive model, and a system cost model as well as an optimization goal of maximizing system utility according to the digital twin-driven edge IoT network slice system model;

[0116] S4. According to the prediction result of the user behavior information of the IoT terminal, and transform the resource allocation problem into a Markov decision process, and use the multi-agent double-delay deep deterministic policy gradient algorithm combined with prioritized experience replay to solve the resource allocation strategy with the optimization goal of maximizing system utility.

[0117] In this embodiment, the physical infrastructure layer in S1 mainly consists of two parts: the access network and the core network. The access network includes base stations, multi-access edge computing servers of the base stations, and IoT terminals. The DT function of the IoT terminals is deployed on the MEC servers of the base stations. By leveraging the advantages of edge computing, the real-time performance and reliability of the DT network are improved. The core network is composed of multiple interconnected servers. Each server runs one or more virtual machines, and each server supports SDN / NFV technology to achieve the orchestration and management of network functions, resources, and traffic data. In addition, some servers in the core network are dedicated to collecting information about the physical network to build the DT of virtual network functions. In the access network, the sets of IoT terminals and base stations are respectively denoted as and Each base station includes an MEC server. Regarding the base station and its MEC server as a whole, the total uplink bandwidth resource of base station m can be expressed as B m . Regarding the core network as a fully connected undirected graph, it can be denoted as G P = {N P , L P}, where N P = {1,..., n,..., N} is the set of physical nodes. The physical nodes are divided into digital twin service nodes and virtual network function service nodes. L P = {l nn′ | n, n′ ∈ N P} represents the set of physical links. The computing resources and storage resources of physical node n are respectively and The bandwidth resource of physical link l nn′ is

[0118] The digital twin network layer monitors the status information of physical infrastructure such as IoT terminals, physical nodes, and physical links in the physical network through MEC servers and digital twin servers, realizes the twin mapping of the physical network, and constructs a digital twin network in a data-driven manner. The digital twin network can not only simulate the behavior of the physical network but also predict potential performance bottlenecks and fault risks, thereby supporting proactive maintenance and optimizing resource allocation. The digital twin network can be represented as an undirected graph G DT = {N DT , L DT}, N DT is the set of DT nodes, including the DT nodes of IoT terminals and the DT nodes of VNF L DTis the set of DT links corresponding to the physical link. Through the DT network, the real-time and historical network status information of IoT terminals is used to predict the future status information of IoT terminals, that is, user behavior information (location / business request), so as to assist the DT network in dynamically adjusting resource allocation. If the resource allocation scheme based on prediction is better than the current one, a new scheme is sent to the physical network for execution through the SDN control layer, which can avoid the negative impact brought by the passive resource allocation scheme and reduce the trial-and-error cost and risk of the actual network.

[0119] The network slicing technology provides service support for the services of IoT terminals. SFC is the core component of network slicing. SFC is composed of a series of VNFs executed in sequence to meet the functional requirements of specific services. A series of SFC sets represent different types of service requests in the network. The s-th SFC can be expressed as where, represents the set of VNFs on is the demand for computing resources, storage resources, and bandwidth resources at time t and represents the set of virtual links. It is assumed that each VNF can only provide services for one type of service, that is, there is The resource demand at time t is the sum of the resources required for all IoT terminal service requests it serves. The service request of IoT terminal u at time t can be expressed as a set of VNFs Introduce the binary variable to represent whether it is associated with , and are the computing, storage, and bandwidth resource demands of respectively, which can be expressed as:

[0120]

[0121]

[0122] Considering that the base station has a fixed coordinate position, the coordinate position of base station m is expressed as (x m , y m ), while the IoT terminal u has dynamic coordinate position changes, which can be expressed as (x u (t), y u (t)). The current movement speed and direction of terminal u are related to the previous state. Therefore, the Gaussian Markov model is used to describe the update of the terminal's coordinate position. Its movement speed and movement direction at time t can be expressed as:

[0123] ​​

[0124] In the formula, and respectively represent the average speed and average moving direction of the terminal u; 0 ≤ σ1′, σ2′ ≤ 1 represents the influence of the previous state of the terminal on the current state; Φ u and Ψ u respectively follow independent Gaussian distributions and

[0125] Then at time t, the coordinate position update process of the terminal u can be expressed as:

[0126]

[0127] In the formula, δ t is the time interval between two adjacent time slots.

[0128] The SA-TCN prediction model established in S2 of this embodiment is:

[0129] SA-TCN characterizes the correlation between different time steps by introducing the self-attention mechanism, and uses the self-attention mechanism to assign different weights to different elements of the time series data. First, the feature vector of the input time series data is Perform positional encoding on it. The method of positional encoding provides the relative or absolute position of the tokens in the input time series data. Adopt the sine positional encoding P e , and the formula is as follows:

[0130]

[0131] In the formula, pos represents the index of the time step; i is the index of the dimension; W H is the total dimension of X I .

[0132] The feature vector after performing the positional encoding operation is X I ′ = X I + P e , input X I ′ into the self-attention layer to calculate the attention scores at each time step, weight different time steps, and learn the relationship between different time steps, so as to capture global dependencies. Through the product of X I ′ and different learnable parameter matrices W Q , W K and W V to obtain the query matrix the key matrix and the value matrix where W h is the dimension of Q and K, and W vLet \(d_V\) be the dimension of \(V\). Attention scores are calculated using \(Q\), \(K\), and \(V\), and the attention scores are normalized using the softmax function. The feature vector after weighting the attention scores is \(X''\). I , and its calculation formula is as follows:

[0133]

[0134] The feature vector \(X''\) after being processed by the self-attention layer I is used as the input of the TCN layer. The TCN will further process the temporal features, capture local temporal dependencies, and increase the receptive field of the model through dilated causal convolutions. At the same time, residual connections are used to avoid the problem of gradient vanishing in long-sequence training. The formula for dilated causal convolutions is as follows:

[0135]

[0136] In the formula, \(F(X''\) I ) is the output of the dilated causal convolution; \(G\) is the size of the convolutional kernel; \(f(g)\) is the \(g\)-th element in the convolutional kernel; \(X''\) t-d·g is the feature vector at different time steps; \(d\) is the dilation factor.

[0137] The output of each layer of convolution is added to the input through a residual connection, and the ReLU activation function is applied to increase non-linearity. The formula is as follows:

[0138]

[0139] To improve the expressive power of the model, the TCN usually adopts a multi-layer convolutional structure. Each layer uses dilated causal convolutions and activation functions to extract features, and the information of each layer is integrated through residual connections. After multiple layers of convolution and residual connection operations, the pooling layer and the fully connected layer are used to map the feature information to the prediction space to obtain the final prediction result.

[0140] The model-agnostic meta-learning method established in S2 of this embodiment is as follows:

[0141] In an IoT environment with limited data volume, traditional user behavior prediction models often have insufficient generalization ability. Especially in a dynamic network environment, the prediction performance will decrease significantly. To enable the prediction model to still maintain high prediction performance in a dynamic network environment, the MAML method is introduced to optimize the generalization ability of the SA-TCN prediction model. The training process of MAML is divided into two stages: inner training and outer optimization. Through meta-training with alternating inner and outer layers on multiple different tasks, the SA-TCN prediction model can learn an optimized initial parameter, so that when facing a new network environment, it can quickly adapt with only a small number of samples and a small number of gradient updates, improving the accuracy and generalization ability of the prediction model.

[0142] Regarding the user behavior prediction task of IoT terminals in different network environments as an independent learning task, that is, the dataset under each network environment corresponds to an independent learning task T i , where represents the input features at time t (such as the historical location of the IoT terminal, the resource requirements of the service request, etc.), represents the true prediction target at the corresponding time (such as the future location of the IoT terminal, the resource requirements of the service request, etc.). To train a prediction model with good generalization ability, MAML adopts a task-level learning strategy, that is, combining multiple tasks into a task set and performing meta-training on these tasks, so that the prediction model can learn how to quickly adapt to different network environment changes. In the meta-training stage, the dataset of task T i is divided into a training set and a test set where the training set is used for inner-layer training to optimize the prediction model parameters for a specific task; the test set is used for outer-layer optimization. After the inner-layer training is completed, it is used to evaluate the generalization ability of the prediction model and calculate the total test loss of all tasks to optimize the initialization parameters of the prediction model.

[0143] In the inner-layer training stage, for task T i , the initialization parameters of the prediction model are θ, and the training set of task T i is used for gradient descent update to obtain the task-specific parameters θ ′. The purpose of inner-layer training is to optimize the prediction performance of the model on the current task T i , that is, to minimize the loss function of the prediction model on the training set, so that the prediction model can better adapt to this specific task. The loss function of inner-layer training can be expressed as: i In the formula, O is the total number of training set samples;

[0144]

[0145] In the formula, O is the total number of training set samples; and respectively represent the input features and the corresponding true prediction target of the o-th sample in the training set of task T i ; represents the prediction result of the prediction model based on the initialization parameters θ on the o-th training set sample.

[0146] The gradient update process of inner-layer training can be expressed as:

[0147]

[0148] In the formula, α intra is the learning rate of task inner-layer training; is the loss function gradient of

[0149] After the inner layer training is completed, MAML enters the outer layer optimization stage to improve the generalization ability of the prediction model. Specifically, MAML uses the test sets of all tasks to calculate the test loss, and optimizes the initial parameters θ of the prediction model by minimizing these test losses, so that the prediction model can share knowledge among different tasks and quickly adapt to new tasks. For task T i , its test loss function can be expressed as:

[0150]

[0151] In the formula, Z is the total number of samples in the test set; and respectively represent the input feature and the corresponding true prediction target of the z-th sample in the test set of task T i ; represents the prediction result of the prediction model based on the specific task parameter θ i ' on the z-th test set sample.

[0152] Furthermore, the total test loss of all tasks is calculated and can be expressed as:

[0153]

[0154] In the formula, I' is the total number of tasks.

[0155] Finally, the initial parameters θ of the prediction model are updated through outer layer optimization, and the update process can be expressed as:

[0156]

[0157] In the formula, β extra is the learning rate of outer layer optimization; is the gradient of the total test loss with respect to the initial parameter θ.

[0158] The end-to-end delay model established in S3 of this embodiment is:

[0159] The delay of the service request can be divided into wireless transmission delay and wired transmission delay. The wireless transmission is the wireless uplink transmission process between the IoT terminal and the MEC. During the wireless transmission process, the coordinate position of the IoT terminal u will affect the signal strength and path loss, thus affecting the signal-to-interference-plus-noise ratio and the wireless transmission rate. Assume that at time t, the IoT terminal u is associated with the MEC server M u (t) closest to it, and is its position is The distance between the MEC server and the terminal u is Assume that the transmission channel of the uplink adopts a fast fading channel model. The transmission power and channel gain of the IoT terminal u are p u (t) and h u (t) respectively. The total interference power of other IoT terminals is I, and the noise power is σ 2 . The path loss factor is θ. Then the signal-to-interference-plus-noise ratio (SINR) of the transmission can be expressed as:

[0160]

[0161] Then at time t, the wireless transmission delay of the uplink of the IoT terminal u can be expressed as:

[0162] R u (t) = B u (t) log2(1 + SINR u (t))

[0163]

[0164] In the formula, B u (t) is the uplink bandwidth allocated by the base station to the IoT terminal u; D u (t) is the size of the transmission data requested by the IoT terminal u for the service; R u (t) is the uplink transmission rate.

[0165] The wired transmission delay is divided into the transmission delay between the MEC and the core network and the transmission delay in the core network. Since the delay of accessing the core network is much higher than the local delay, the transmission delay between the MEC and the core network is set to a fixed value T MC . The transmission in the core network is the transmission of service requests between servers, and it can be regarded as an SFC. That is, the wired transmission delay of the service request can be divided into the processing delay of the SFC at the physical node and the communication delay between the physical links. Introduce a binary variable to represent the mapping relationship between the VNF node and the server, represents is mapped to the physical node n, otherwise Similarly, represents the mapping relationship between the virtual link and the physical link. Assume is the service for the business request of the IoT terminal u. Then its processing delay at the physical node can be expressed as:

[0166]

[0167] In the formula, represents on processing rate; represents the processing rate coefficient of service request data.

[0168] At time t, its communication delay can be expressed as:

[0169]

[0170] In the formula, represents the physical distance of physical link l nn′ ; c represents the speed of light.

[0171] Due to the dynamic changes of service requests, the SFC needs to be dynamically scheduled to ensure service quality. The scheduling of the SFC is reflected in the possible migration of VNFs. During the migration of VNFs, a certain migration delay will be generated, which directly affects the quality of the entire service request. At time t, the total migration delay can be expressed as:

[0172]

[0173] In the formula, is the size of the migration data; is the delay required to transmit a unit of data along the migration path; Ω is a positive coefficient; is the network routing hop distance between server nodes.

[0174] The total delay of wired transmission is Then the end-to-end delay can be expressed as:

[0175]

[0176] The DT credibility model established by S3 in this embodiment is:

[0177] Due to the complex and changeable network environment of mobile IoT terminals, deploying DT on MEC servers can not only bring low-latency computing and efficient resource use, but also effectively reduce synchronization delay and improve the adaptability and accuracy of DT in a time-varying network environment. Similarly, reducing synchronization delay enables DT to quickly update the model and react after receiving the latest status of physical devices, thereby improving the overall efficiency, real-time performance and credibility of the system. Since the DT of VNF is generally deployed in core network servers connected by wire, its network environment is relatively stable and the delay is low. The DT synchronization delay and credibility problems in the core network are relatively small and less affected by the wireless network environment. Therefore, the DT credibility of IoT terminals on the wireless side is mainly considered.

[0178] There are mainly two factors affecting the DT credibility of IoT terminals, namely DT mapping deviation and DT synchronization delay deviation. For DT mapping deviation, the impacts of computing power (i.e., CPU frequency) and storage capacity are mainly considered, and the CPU frequency deviation Δf c (t) and storage capacity deviation Δf m (t) are introduced to comprehensively measure the DT mapping deviation of IoT terminals. Then, the DT mapping deviation can be expressed as:

[0179] DT map =ω c Δf c (t)+ω m Δf m (t)

[0180] In the formula, ω c and ω m are the weight parameters of CPU frequency deviation and storage capacity deviation.

[0181] The DT of IoT terminal u is deployed on the associated MEC server M u (t). For DT synchronization delay deviation, the impacts of wireless transmission delay and MEC server processing delay are mainly considered. The wireless transmission delay can be obtained from the above text as The processing delay refers to the time when the MEC server calculates and updates the DT model after receiving the synchronization data of the IoT terminal. This is related to the computing resources of the MEC server M u (t) and can be expressed as: In the formula,

[0182]

[0183] In the formula, is the size of DT synchronization data; is the data packet processing coefficient of the MEC server.

[0184] Then, the DT synchronization delay deviation can be expressed as:

[0185]

[0186] In the formula, ω wire and ω deal are the weight parameters of wireless transmission delay and processing delay.

[0187] Then, for IoT terminal u, its DT credibility can be expressed as:

[0188]

[0189] In the formula, ψ DT∈(0,1], the closer its value is to 1, the higher the credibility of DT; ω map and ω syn are weight parameters, which determine the influence degree of mapping deviation and synchronization delay deviation on credibility.

[0190] The penalty / incentive model established in S3 in this embodiment is:

[0191] The reasonable control of resource utilization is crucial for system performance and stability. Excessive resource utilization may lead to resource overload, increase latency and even cause system crashes; while too low utilization will result in resource waste and increased operating costs. Since the resources of the MEC server are limited, in order to optimize network load and meet more user service requests, it is necessary to introduce a penalty / incentive mechanism based on resource utilization to analyze and improve the load balancing performance of the network.

[0192] At time t, the computing resource utilization rate storage resource utilization rate and bandwidth resource utilization rate on physical node n can be respectively expressed as:

[0193]

[0194]

[0195] Let ω1, ω2 and ω3 be the influence weights of computing resources, storage resources and bandwidth resources on resource utilization rate, and ω1 + ω2 + ω3 = 1. Then for the resource utilization rate η of physical node n n can be expressed as:

[0196]

[0197] Let be the lower and upper limits of the threshold of the resource utilization rate of the physical node. If its resource utilization rate exceeds the upper threshold it means that the node is in an overloaded state. If its resource utilization rate is less than it means that the node is in a lightly loaded state. Both overload and light load will affect the overall network performance, and penalties are given to both the overloaded and lightly loaded parts. The greater the difference from the upper and lower threshold values, the more penalties are imposed. Assume that the unit overload / light load penalty is π. And a fixed incentive is given for the case of balanced resource utilization rate Then for physical node n, its penalty function e n (t) can be expressed as:

[0198]

[0199] Then at time t, the total penalty E of the systemtotal (t) can be expressed as:

[0200]

[0201] The system cost model established in S3 of this embodiment is:

[0202] To ensure the real-time update of the mobile IoT terminal DT, the DT will migrate as the IoT terminal associates with a new base station, that is, the DT of the IoT terminal needs to be migrated to a new MEC server. When the DT migrates from the source MEC server to the target MEC server, the DT historical data and model parameters will be migrated through a wired fiber optic cable, thus generating the corresponding DT migration cost. At time t, if the DT of IoT terminal u migrates from MEC server m to MEC server m′, then its DT migration cost can be expressed as:

[0203]

[0204] In the formula, κ DT is the DT migration cost factor; is the size of the migrated data of the DT; is the physical distance between MEC servers m and m′.

[0205] The mobility of the IoT terminal directly affects the DT migration. The DT migration occurs when M u (t - 1) ≠ M u (t). In this case, at time t, the total DT migration cost of the system can be expressed as:

[0206]

[0207] When the core network server is overloaded or fails, etc., it is necessary to migrate the VNF instances carried by the physical nodes to new physical nodes to maintain network load balancing and service quality. At time t, if migrates from physical node n to physical node n′, then the VNF migration cost can be expressed as:

[0208]

[0209] In the formula, κ VNF is the VNF migration cost factor.

[0210] At time t, the total VNF migration cost of the system can be expressed as:

[0211]

[0212] Then at time t, the total cost C of the systemtotal It can be expressed as:

[0213] C total = C DTmig (t) + C VNFmig (t)

[0214] The optimization objective established in S3 of this embodiment is:

[0215] The resource allocation problem of the system can be expressed as: how to minimize the system cost and the penalty for unbalanced network resource utilization under the constraints of end-to-end delay and DT credibility. To unify the units of cost and penalty, they are normalized, and the utility function is designed as:

[0216]

[0217] In the formula, represents the maximum penalty suffered by the network due to unbalanced resource utilization; represents the maximum migration cost of the system; β1 and β2 represent the importance of cost and penalty to the utility function, and β1 + β2 = 1.

[0218] Then, the optimization objective can be expressed as:

[0219]

[0220] In the formula, C1 means that a VNF can only be mapped to one physical node; C2 means that a virtual link is only mapped to one physical link; C3 means that an IoT terminal is only associated with the nearest MEC server; C4 is the end-to-end delay constraint, τ max is the end-to-end maximum delay constraint; C5 represents the DT credibility constraint of the IoT terminal, ψ max is the maximum credibility constraint; C6 - C8 mean that the computing and storage resource requirements of the VNF must be within the resource upper limit of the physical node, and the bandwidth resource requirements of the virtual link cannot exceed the bearing limit of the physical link; C9 is a binary variable for the association relationship between the VNF node and the IoT terminal; C10 - C11 are binary variables for the mapping relationship between the VNF node and the virtual link.

[0221] The Markov decision process in S4 of this embodiment is modeled as:

[0222] To transform the resource allocation problem into an MDP, a network slice is regarded as an agent. Suppose there are K agents, and the agents maximize the utility of the system by optimizing the allocation of resources such as base stations and servers. The key elements of the MDP model include the state space, action space, and reward function, which can be defined as where

[0223] (1) State space: The state space represents the configuration state of the system at a certain moment, including the allocation of resources and other information of the system. The state of agent k at time t is defined as where represents the resource configuration state of the base station; and are the sets of computing resource requirements and storage resource requirements of the VNF; represents the set of bandwidth resource requirements of the virtual link; represents the connection state of the IoT terminal, including the location of the terminal and the relationship with the MEC server.

[0224] (2) Action space: The action space represents the set of actions that the agent may choose in the current state. The action space of agent k at time t can be defined as where a band is the wireless bandwidth resource allocation, which determines the wireless bandwidth resources allocated by the base station to each IoT terminal to ensure communication quality and meet the end-to-end delay constraint; a vmap is the set of VNF node mappings, which determines the deployment location of each VNF node on the server; a lmap is the set of virtual link mappings, which determines the mapping location of the virtual link on the physical link.

[0225] (3) Reward function: The reward function is the feedback signal obtained by the agent from the environment after executing an action, which is used to measure the effect of the current decision and guide the policy update. To encourage the agent to learn the optimal resource allocation strategy, at time t, agent k in state taking action will obtain an instantaneous reward The instantaneous reward of the agent is defined according to the contribution degree of the agent to the system utility, which can be expressed as:

[0226]

[0227] In the formula, H m,k is a binary variable recording the association information between base station m and agent k; U′ m represents the contribution degree of base station m to the system utility; is for recording and the binary variable of the association information of agent k; represents 's contribution degree to the system utility. Then at time t, the global reward of the system can be expressed as:

[0228]

[0229] The modeling process of the PER-MATD3 algorithm in S4 of this embodiment is as follows:

[0230] The basic framework of PER-MATD3 is the Actor-Critic structure. Each agent k includes an Actor network with parameter λ k and two Critic networks with parameters μ k,1 and μ k,2 , as well as their corresponding target networks. The target network parameters are λ′ k , μ′ k,1 , μ′ k,2 . The Actor network is mainly used to learn the policy, that is, to select the optimal action according to the current state and optimize the policy π by interacting with the environment. The Critic network is responsible for evaluating the value of the action selected by the Actor network, that is, calculating the Q value of the state-action pair and guiding the Actor network to update the policy. k .

[0231] The PER-MATD3 algorithm uses two Critic networks and takes the minimum Q value when calculating the TD target value to reduce the bias of Q value overestimation and avoid overfitting problems in the policy learning process. The Prioritized Experience Replay (PER) mechanism can perform priority sorting according to the importance of each sample, sample high-priority samples more frequently from the experience pool, thereby improving the learning efficiency and accelerating the convergence speed. For agent k, the priority Λ t of the i-th sample {S t , A t , R t+1} in the experience replay pool is related to the TD error and can be expressed as: i In the formula, γ represents the discount factor;

[0232]

[0233] represents the Q value calculated by the target Critic network, and A is determined by the target Actor network; t+1 is a very small positive constant to ensure that samples can still be sampled even when the TD error is zero. Samples with larger TD errors are given higher priorities, and conversely, samples with smaller TD errors are given lower priorities. To enhance the exploration ability, smooth regularization is added to the target policy, that is, noise is added to the action selected by the target Actor network. Then the target action A

[0234] t+1 ​It can be expressed as:

[0235]

[0236] In the formula, is the policy function of the target Actor network; is the noise that follows a Gaussian distribution, and the random noise parameter is c′ to achieve a smoother Q-value estimation.

[0237] Samples in the experience pool are probabilistically sampled according to the priority Λ i Then the sampling probability of the i-th sample is:

[0238]

[0239] In the formula, Γ represents the influence degree of the priority on the sampling probability; is the total number of samples in the experience replay pool.

[0240] Since the PER mechanism destroys the uniform sampling distribution of the original experience pool and may introduce biases, the importance sampling weight Θ i is used for correction and can be expressed as:

[0241]

[0242] In the formula, is a hyperparameter used to control the degree of correction. When it represents traditional experience replay.

[0243] During the training process, for agent k, a small batch of quadruples is randomly sampled from the experience replay pool D as sample data and input into the Actor-Critic network for optimization. The goal of the Critic network is to minimize the TD error, and the Critic network objective loss function can be expressed as:

[0244]

[0245] In the formula, represents the TD target value, and there is

[0246] Then, the parameters of the Critic network are updated using the gradient descent method:

[0247]

[0248] In the formula, β critic is the learning rate of the Critic network.

[0249] The goal of the Actor network is to learn the optimal policy to maximize the expected cumulative discounted reward Since the agent cannot directly optimize the future cumulative reward, the Actor network approximately optimizes the policy by maximizing the Q-value evaluated by the Critic network, so that the actions it selects can obtain the maximum Q-value. The objective loss function of the Actor network can be expressed as:

[0250]

[0251] Then, the gradient ascent method is used to update the parameters of the Actor network:

[0252]

[0253] where β actor is the learning rate of the Actor network. Since the Actor network depends on the Q-value evaluated by the Critic network, the Actor network will only be updated after the Critic network is updated and after a fixed number of steps T up .

[0254] The parameters of the target network are updated in a soft update manner. The parameter update processes of the target Actor network and the target Critic network can be expressed as:

[0255] λ′ k ← τλ k +(1 - τ)λ′ k

[0256] μ′ k,j ← τμ k,j +(1 - τ)μ′ k,j j = 1, 2

[0257] where τ is the update rate of the target network.

[0258] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A digital twin-driven IoT user behavior prediction and resource allocation method, characterized by: The following steps are involved: S1: Build a digital twin-driven edge IoT network slicing system model, including the physical infrastructure layer, digital twin network layer, and control layer; S2: Establish a SA-TCN time series prediction model to predict user behavior information. The SA-TCN model uses a self-attention mechanism to enhance the capture of the dependency relationship between time series data features at different time steps, and combines a model-independent meta-learning method to improve the generalization ability of the prediction model; S3: Based on the system model, an end-to-end delay model, a DT credibility model, a penalty / incentive model and a system cost model are established to construct a mathematical model with maximizing system utility as the optimization goal; S4: Based on the predicted user behavior information, the resource allocation problem is transformed into a Markov decision process, and the multi-agent double-delay deep deterministic policy gradient algorithm PER-MATD3 combined with priority experience replay is used to solve the resource allocation strategy.

2. The method for predicting user behavior and allocating resources for the Internet of Things driven by digital twins according to claim 1 is characterized in that: In S1, the physical infrastructure layer includes an access network and a core network; the access network includes a base station, a MEC server of the base station and a set of IoT terminals; the core network is a fully connected undirected graph, consisting of multiple servers, each of which supports SDN / NFV technology; the digital twin network layer builds a virtual mapping by synchronizing the status of physical devices in real time, and dynamically adjusts the resource allocation plan based on the predicted IoT terminal location and business needs; the control layer sends the resource allocation strategy to the physical network through the SDN controller.

3. The method for predicting user behavior and allocating resources for the Internet of Things driven by digital twins according to claim 1 is characterized in that: In S2, the SA-TCN model includes: Self-attention layer: performs sinusoidal position encoding on the input time series data and calculates the attention score through the query matrix Q, key matrix K and value matrix V; TCN layer: uses multi-layer dilated causal convolution to extract temporal features, integrates the outputs of each layer through residual connections, and applies the ReLU activation function; Meta-learning method: Model-independent meta-learning (MAML) is used for training. The inner-layer training optimizes the task-specific parameters, and the outer-layer optimization updates the model initialization parameters to improve the generalization ability.

4. The method for predicting user behavior and allocating resources for the Internet of Things driven by digital twins according to claim 1, characterized in that: The attention score of the self-attention layer is calculated as: Among them, d k To determine the dimensions of the query matrix and key matrix, Q, K, and V are obtained by multiplying the input time series data with the learnable parameter matrix, respectively.

5. The method for predicting user behavior and allocating resources for the Internet of Things driven by digital twins according to claim 1, characterized in that: In S3, the end-to-end delay model includes wireless transmission delay, wired transmission delay and VNF migration delay, and the calculation formula is: in, is the uplink transmission delay of IoT terminal u, For core network processing and communication delays, Delay for VNF migration.

6. The method for predicting user behavior and allocating resources for the Internet of Things driven by digital twins according to claim 1, characterized in that: In S3, the DT credibility model is: in, is the DT mapping deviation, is the DT synchronization delay deviation, α1 and α2 are weight parameters.

7. The method for predicting user behavior and allocating resources for the Internet of Things driven by digital twins according to claim 1, characterized in that: In S3, the optimization goal is to maximize the system utility: Among them, η1 and η2 are weight coefficients, Penalty(t) is the resource utilization penalty, and Cost(t) is the system migration cost.

8. The method for predicting user behavior and allocating resources for the Internet of Things driven by digital twins according to claim 1, characterized in that: In S4, the resource allocation problem is modeled as a Markov decision process MDP, whose state space includes base station resource allocation state, VNF demand set and IoT terminal connection state; the action space includes wireless bandwidth allocation, VNF node mapping and virtual link mapping; the reward function is designed based on the system utility contribution.

9. The method for predicting user behavior and allocating resources for the Internet of Things driven by digital twins according to claim 1, characterized in that: The PER-MATD3 algorithm includes: Prioritized Experience Replay PER: Sample priority is calculated based on TD error, and the sampling probability is: Among them, Λ i represents the priority of sample i; Γ represents the influence of priority on sampling probability; is the total number of samples in the experience replay pool; Dual Critic Network: Two Critic networks are used to calculate the Q value, and the minimum value is taken as the TD target to suppress overestimation; Target network soft update: The target network parameters are updated through λ k ′←τλ k +(1-τ)λ k ′ is updated, where τ is the update rate.

10. The method for predicting user behavior and allocating resources for the Internet of Things driven by digital twins according to claim 9, characterized in that: The objective loss function of the Critic network is: represents the TD target value, and has Where γ is the discount factor, is the Q value estimate of the target Critic network; Θ i represents the importance sampling weight.

Citation Information

Cited By

  • Edge gateway monitoring data transmission method and system combined with neural network

    CN120455498A