Distributed network slice switching method based on multi-agent deep circulation double-Q network

By employing a multi-agent deep recurrent dual-Q network algorithm, combined with parameterized noise and gated recurrent units, the problems of high communication load and low action exploration efficiency in heterogeneous cellular networks are solved, achieving efficient and stable network handover decision-making, and making it suitable for distributed handover in heterogeneous cellular networks.

CN121240154APending Publication Date: 2025-12-30YANCHENG INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511357587.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

In heterogeneous cellular networks, existing network handover strategies based on deep reinforcement learning suffer from problems such as high communication load, incomplete state observation, and low action exploration efficiency, resulting in handover decision delays and instability, making it difficult to meet real-time service quality requirements.

Method used

A multi-agent deep recurrent double-Q network algorithm with centralized training and distributed execution is adopted. By introducing parameterized noise and gated recurrent units, the exploration efficiency is improved and distributed local decision-making is achieved by utilizing historical observation data.

Benefits of technology

It reduces communication overhead, improves the accuracy and stability of handover decisions, enhances the system's real-time response capability and scalability, and is suitable for network handover in high-dimensional action spaces and dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121240154A_ABST
    Figure CN121240154A_ABST
Patent Text Reader

Abstract

The invention provides a distributed network slice switching method based on a multi-agent deep cycle double-Q network. The method comprises the following steps: constructing a heterogeneous cellular network scene based on end-to-end network slices; defining a service type of the network slice and establishing a service model and a switching model required by a switching problem; modeling a switching problem into a distributed partially observable Markov decision process model, and intelligently selecting a proper network slice in the network scene according to an actual demand and a network condition of terminal equipment so as to meet a network service demand of the terminal equipment; the distributed Markov decision process model is optimized and solved by adopting a multi-agent deep cycle double-Q network algorithm to obtain a switching strategy, and then slice distributed switching is performed on the switching system based on the end-to-end network slices based on the switching strategy, so that the defects of an existing switching algorithm are overcome; and an effective solution is provided for switching of the network slices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication, and specifically relates to a distributed network slice switching method based on a multi-agent deep cyclic dual-Q network. Background Technology

[0002] The purpose of network slicing technology is to provide diversified network services to meet the needs of businesses with different resource requirements and performance. Although network slicing offers promising services, switching between network slices faces complex problems and significant challenges in practical applications.

[0003] With the rapid development of communication technologies and the rise of emerging applications such as the Internet of Things, intelligent transportation systems, and smart cities, the demand for network services is increasing rapidly, especially in terms of data rates, coverage, and low latency, which places higher demands on network resource management and service quality. As one of the core technologies of 5G, network slicing divides the physical network into multiple virtual networks through virtualization technology, providing customized services for different applications and improving network flexibility and efficiency.

[0004] Heterogeneous cellular networks are a crucial technology for 5G and the future 6G era, offering high capacity, high data rates, low latency, and wide coverage, particularly enhancing the communication experience for edge terminals. Heterogeneous cellular networks deploy small cells beneath macro base stations and combine them with various wireless access technologies to meet the demands of next-generation wireless communication systems. In environments with dense buildings and numerous obstacles, the deployment of small cells is critical; its core concept is to bring the network closer to the terminal, thereby improving connection quality and data transmission rates. Based on coverage area and terminal capacity, small cells can be categorized into picocells and femtocells, among others. Through a multi-layered base station architecture combined with various wireless access technologies, heterogeneous cellular networks ensure stable, high-speed communication in complex environments, laying a solid foundation for the 5G and future 6G era.

[0005] Deploying end-to-end network slicing in heterogeneous cellular networks can significantly improve data rates and network capacity. However, the dense deployment of base stations leads to frequent handovers and high handover overhead, posing a major challenge to network performance. While end-to-end network slicing enhances the flexibility and performance of mobile communication networks, such as capacity, latency, and transmission rates, it also introduces more management challenges. Particularly in mobility management, the service range of a network slice is limited by the coverage area of ​​the base stations it relies on. In environments with dense base station deployments, frequent terminal movement can lead to frequent handovers, increasing the risk of radio link failures and impacting service quality. Furthermore, since different slices are only associated with a subset of base stations in a specific area, relying solely on traditional handover strategies such as signal strength and load may result in selecting base stations that do not support the target service, leading to handover failures. Core network resource constraints are also a factor to consider. Therefore, a thorough study of handover mechanisms is needed, taking into account the characteristics of network slicing.

[0006] In current technologies, the most common network handover strategies based on deep reinforcement learning generally employ a centralized training and execution architecture. While this approach can theoretically achieve a globally optimal strategy, in real-world heterogeneous networks, the large number of terminals and dense base station deployments cause the network state space to expand rapidly, resulting in the training process handling an extremely large amount of state information. Furthermore, in centralized systems, all network nodes need to periodically transmit state information back to the central controller and receive decision instructions. This communication mode introduces a significant signaling load in practical deployments, especially in scenarios with frequent mobile terminal handovers and drastic dynamic changes in network state. This leads to a significant increase in decision latency, weakening the timeliness and reliability of handover decisions and failing to meet the requirements for real-time service quality assurance. Existing distributed handover schemes generally fail to fully utilize historical observation data over time, making it difficult for the agent to accurately recover or predict the environmental state when faced with missing information. This affects the accuracy and continuity of handover judgments, making the strategy susceptible to noise interference or large fluctuations, resulting in insufficient stability. Moreover, in reinforcement learning, a fundamental and crucial issue is the balance between exploration and utilization. Especially when the state space of the environment is large, or when there are deceptive or sparse rewards, efficient exploration becomes a significant challenge. Existing deep reinforcement learning-based switching methods often employ an ε-greedy strategy for action exploration, in which the agent selects a random action with a low probability. This method explores by adding noise to the agent's actions, but it is inefficient.

[0007] This invention proposes a distributed network slice switching algorithm based on a multi-agent deep recurrent double-Q network, aiming to solve the problems of excessively high state dimension, heavy communication load, incomplete state observation, and low action exploration efficiency in existing methods. This method abstracts the distributed network slice switching problem into a distributed locally observable Markov decision process model to more realistically characterize the decision-making characteristics of multi-agents in dynamic network environments. To balance the global optimality of the policy with real-time execution, a learning framework of centralized training and distributed execution is adopted. Based on this, parameterized noise is added to the connection weights of the neural network to improve the algorithm's exploration efficiency. Furthermore, a gated recurrent unit is introduced to extract and memorize the historical observation and behavioral sequence information of the agents during the interaction process, thereby enhancing the state estimation capability. This allows each agent to make more accurate and stable switching decisions based solely on its own local observations and memorized information during actual execution. Summary of the Invention

[0008] One of the objectives of this invention is to provide a distributed network slicing switching method based on a multi-agent deep recurrent dual-Q network, which solves the problems of high communication overhead and decision complexity in centralized switching algorithms and reduces signaling overhead during communication; it solves the local observability problem in multi-agent environments and improves the utilization rate of historical observation data; and it solves the problem of low efficiency in existing action exploration strategies and improves the exploration capabilities of agents.

[0009] This invention provides a distributed network slicing switching method based on a multi-agent deep recurrent dual-Q network, comprising:

[0010] S1: Construct a heterogeneous cellular network scenario based on end-to-end network slicing;

[0011] S2: Define the service types of network slices and establish the service model and switching model required for network slice switching issues;

[0012] S3: The handover problem is modeled as a distributed partially observable Markov decision process model. In the heterogeneous cellular network based on end-to-end network slicing, the slice handover intelligently selects the appropriate network slice according to the actual needs of the terminal device and the network conditions to meet the different network service needs of the terminal device.

[0013] S4: The distributed Markov decision process model is optimized and solved using a multi-agent deep recurrent double-Q network algorithm to obtain a distributed handover strategy. Then, based on the distributed handover strategy, slice distributed handover is performed on the heterogeneous cellular network system based on end-to-end network slices.

[0014] Furthermore, in S1, the constructed network scenario includes: a wireless access network, a core network, and a bearer network;

[0015] The wireless access network includes: terminal devices, macro base stations, small base stations, home base stations, and network slices. Different types of network slices are deployed on one or more base stations. One or more types of network slices are deployed on a base station. Terminal devices connect to the network slices on various types of base stations through the wireless access network to obtain relevant network services.

[0016] The core network mainly includes Network Function Virtualization (NFV) and Software-Defined Networking (SDN). NFV reduces network costs and improves network flexibility, while SDN enables dynamic configuration and management of network functions.

[0017] The resource status of the bearer network and the core network is abstracted as the available bandwidth of the core network.

[0018] Furthermore, in S1, end-to-end network slices have been deployed during the network initialization phase according to predefined SLAs for different services, and the association between base stations and network slices remains unchanged during the simulation. Regarding user requirements, two parameters are used to describe the quality of service requirements of terminal devices. These parameters include:

[0019] Represents network slice n i The minimum threshold for the guaranteed transmission rate;

[0020] T i , representing network slice n i It can guarantee the maximum time during which the transmission rate is below the minimum threshold.

[0021] To serve as many end devices as possible, it is assumed that each network slice provides services to end devices using the minimum transmission rate required by the service.

[0022] Furthermore, in S2, the service types of network slices are defined, and the task model and handover model required for the network slice handover problem are established, specifically including:

[0023] The task model employs a fully buffered service mode, meaning each terminal device (UE) in the network always has its own data stream to be transmitted. To ensure consistency with the quality of service provided by network slicing, two parameters are considered to describe the UE's quality of service requirements from the user's perspective. These parameters include:

[0024] The minimum threshold representing the transmission rate;

[0025] τ j This indicates the tolerable time, i.e., the terminal device u. j The maximum time during which its transmission rate is allowed to be below the minimum threshold.

[0026] To serve as many terminal devices as possible, it is assumed that each network slice provides service to terminal devices using the minimum transmission rate required by the service. Therefore, at time step t, the number of terminal devices u... j Access BS-NS to h l The estimated wireless transmission network bandwidth consumed is calculated as follows:

[0027]

[0028] Furthermore, the terminal device u was defined using Shannon's theory. j via base station b k Access BS-NS to h l Transmission rate at time and Calculated as:

[0029]

[0030] Among them, W B This represents the bandwidth obtained by each user. At time step t, the terminal device u j With base station b k The signal interference plus noise ratio between them. Furthermore, Calculated as:

[0031]

[0032] Where, p i Indicates base station b i The transmission power, G i,j (t) represents the terminal device u j and base station b k The wireless channel gain between For other base stations to terminal equipment u j The interference is N0, which is Gaussian white noise.

[0033] For ease of description, at time step t, respectively using and To represent when the terminal device u j Access BS-NS to h l At the same time, ensure the core network bandwidth and radio access network bandwidth required to meet the user service quality requirements of its mission.

[0034] The handover model is as follows: For ease of description, it is assumed that all terminal devices must make a handover decision at every time point. However, if the base station-network slice pair connected at the previous time point was selected, the actual handover process will not initiate. This is due to the random movement of terminal devices, especially when channel quality deteriorates. In this case, terminal device u... j It will not consume BS-NS pairs for h l The resources will increase latency by 1. When the terminal device u j Continuous delay, and:

[0035]

[0036] Then the terminal device u j An interruption occurs at time step t. Where γ j (t0) represents the terminal device u at time t0. j The obtained real-time transmission rate. Furthermore, if the terminal device triggers a handover process within the current time step, it will be moved to the end of the UE queue.

[0037] Furthermore, in S3, the switching problem is modeled as a distributed partially observable Markov decision process model, specifically including:

[0038] The network slicing handover problem is modeled as a distributed partially observable Markov decision process, with the optimization objective being to maximize the average cumulative reward associated with the UE's service availability, handover cost, and interruption penalty. The distributed partially observable Markov decision process model can be represented mathematically by a tuple (I, S, A, P, R, O, γ), with detailed definitions of each element as follows:

[0039] 1) I is the set of finite intelligent agents with indices from 1 to n, which is the set of all terminal devices in the system.

[0040] 2) S is a finite set of states, that is, the set of states that exist in the switching model system.

[0041] 3) O is the joint local observation set, where o j ∈O is the terminal u j Local observations.

[0042] 4) A is a finite set of joint actions. At each time step, terminal u j Based on local observations j Perform action a j a = (a1,...,a) U )∈A is a set of joint actions.

[0043] 5) P is the state transition probability, and the state of the environment is determined by the transition function: S×A×S→[0,1).

[0044] 6) R is the reward function, and S×A→R represents the reward that the agent receives after performing an action in the current state.

[0045] 7) γ is a discount factor, and γ∈[0,1].

[0046] Furthermore, the detailed definitions of the specific states, actions, and rewards in S3 are as follows.

[0047] The global state includes both the network state and the local state; therefore, at time step t, the global state is defined as s. G,t =[s N,t ,s U,t ]∈S, where s N,t and s U,t They are defined as follows:

[0048]

[0049] s U,t =[s 1,t ,...,s j,t ,...,s U,t ]

[0050] Among them, s j,t For terminal k j The local state is specifically defined as:

[0051]

[0052] Among them I j For terminal u j When making a decision, the ID, d, of the base station-slice pair accessed in the current time step. j Then the terminal u was recorded. j The remaining available time before service interruption. Because the association configuration between the base station and network slice (BS-NS) remains static for a specific time period, the state vector dimensions of all user equipment (UE) within this area are completely consistent. Based on this, terminal u j Local observations are defined as o j,t =[s N,t ;s j,t ].

[0053] At each decision time t, the terminal device needs to select one from the available base station-slice combinations for access. This action is defined as the terminal device's selection behavior of a base station-slice pair. Therefore, A = {h} l} l=1,2…K′×S , where h lThis is the ID of the base station-slice pair. Each terminal device needs to select a BS-NS pair for access at time step t. Terminal u j At time step t, the action a is defined as: a j,t =h l .

[0054] Due to the stability of the base station-slice association, the set of available actions for all terminals within the region is exactly the same. Therefore, the joint action at time t can be represented as a tuple of all terminal actions: a t =(a 1,t ,...,a j,t ,...,a U,t ).

[0055] To define terminal u j The reward function is designed with three types of utility: the utility of the terminal device being served (R1), the handover cost (R2), and the interruption penalty (R3). Given the differences in decision-making among terminal devices during the handover process, the current decision of terminal device u... j The reward can be represented as:

[0056]

[0057] Therefore, the optimization objective of the switching decision process is the cumulative reward obtained by all terminal devices at each time step, which is expressed as:

[0058]

[0059] Furthermore, in step S4, parameterized noise is added to the fully connected layer of the neural network. The original fully connected layer only needs to learn the weights w; in this invention, the noise parameter ε is added. w and ε b This noise is added to the connection weights of the fully connected layer. In this way, the fully connected layer not only needs to learn the mean μ of the weights w, but also their variance σ, and both the mean and variance are learned as network parameters. Introducing noise into the fully connected layer randomizes the output Q-function, thereby improving the agent's exploration ability. Gaussian noise decomposition is used to generate noise ε; it is assumed that the fully connected layer generates p independent Gaussian noise ε based on the number of neurons. i The next layer generates q independent Gaussian noise ε based on the number of neurons. j Therefore, p+q noise particles will eventually be generated in these two layers. Each and It can be represented as:

[0060]

[0061] Where f is a function, used in this invention

[0062] Furthermore, in S4, in actual heterogeneous networks, since agents cannot obtain the global state of the system, they can usually only make judgments based on partial observation information from the environment. To alleviate the resulting partial observability problem, a recurrent neural network is introduced, enabling agents to memorize historical observation data and actions, thereby comprehensively considering past information and improving the accuracy of decision-making.

[0063] Like LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit) is an improved structure of traditional Recurrent Neural Networks (RNNs), primarily used to address the vanishing gradient problem that RNNs often encounter when handling long-term dependencies. While their actual performance is similar, GRU, due to its simpler structure, is often more efficient when processing large-scale data. Specifically, GRU updates the hidden state by merging the input and forget gates of LSTM into a single update gate, controlling the retention of past hidden states at the current time step. Furthermore, GRU introduces a reset gate, which combines the memory units and hidden layer functionality of LSTM to determine how much state information from the previous time step should be retained when calculating the current candidate hidden state. This design not only simplifies the model structure but also reduces the number of parameters, thereby improving training speed and computational efficiency.

[0064] GRU maintains the same input-output structure as traditional RNNs, with the input s at each time point being... t The state h of the previous moment t-1 They work together in the current calculation process, where h t-1 It carries information from previous points in time. GRU controls the flow of information by introducing update and reset gates, enabling the network to retain or discard information at different time steps. The principle behind its structure is as follows:

[0065] z t Represents the update gate, which is used to control the hidden state h of the previous time step. t-1 To what extent should it be preserved, and what is the current candidate hidden state h? t To what extent should it affect the hidden state h at the current moment? t The calculation method for the update gate is defined as follows:

[0066] z t =σ(W sz s t +W hz h t-1 +b z )

[0067] In the formula, σ represents the Sigmoid activation function, and W represents the weights.

[0068] r t Represents resetting the gate, determined by the current input s. t and the previous hidden state h t-1 The decision mechanism adjusts the degree to which past hidden state information influences the current state. If the current input is unrelated to previous information, r... t It can then take effect to remove h t-1 The impact. The calculation method for resetting the door is defined as follows:

[0069] r t =σ(W sr s t +W hr h t-1 )

[0070] Hidden state h t It is z t h t-1 and h t It is obtained by performing calculations on z. t When h approaches 1 t-1 Approximately equal to h t The new hidden state almost completely preserves the h from the previous moment. t-1 Information reflects a high degree of dependence on past information; conversely, when z t When the value is close to 0, the new state is mainly determined by the current input s. t The decision reduced the number of previously hidden states (h). t-1 Due to the influence of this, this flexible information processing mechanism allows GRU to capture both long-term dependencies and adapt to short-term changes. The calculation method for hidden states is defined as follows:

[0071] h t =(1-z)h t ′+z t ⊙h t-1

[0072] Candidate hidden state h t By resetting the gate's output r t and the hidden state h from the previous moment t-1 Perform element-wise multiplication, and the current input s t The combination is processed by the Tanh function to obtain h. t The range of ' is between -1 and 1. When it is close to 0, the result of element-wise multiplication is close to 0, indicating that information from the previous state is almost completely ignored, allowing GRU to focus on the current input and adapt to capturing short-term dependencies; conversely, when r is close to 1, the result of element-wise multiplication is close to 0. t When it approaches 1, then h t 'Preserves past state h't-1 A large amount of information is beneficial for maintaining long-term memory. The calculation method for candidate hidden states is defined as follows:

[0073] h t =Tanh(W) s s t +W h (r t ⊙h t-1 )+b n )

[0074] In the MA-DDGQN algorithm model, a GRU neural network is added before the fully connected layer of the Q-value network based on the DQN algorithm, and then based on historical information h t and current state s t To predict the Q value, h t This is the input returned by the network at the previous time step. The MA-DDGQN algorithm then uses a Q-value network to calculate the Q-value of the action and obtains the optimal action corresponding to the maximum Q-value. This optimal action is then fed into the target value network for Q-value calculation. In other words, the optimal action is first found using the Q-value network, as shown in the following formula:

[0075]

[0076] Among them, h t - The hidden state, θ, represents the output of the next state after processing by the GRU neural network and serves as the input to the fully connected layer. t Q represents the Q-value network parameter, Q(h) t - Let ,a;θ) represent a series of Q-values ​​for the next state. The significance of this formula is to derive the optimal action corresponding to the next state. Then, using this selected action a... max (h t - ,θ t ) Calculate the target Q-value within the target network, i.e.:

[0077] y=r+γQ′(h t - ,a max (h - ,θ t );θ t - )

[0078] Where y represents the target Q value, r represents the reward from environmental feedback, γ is the discount factor, and θ t - These are the parameters of the target value network. Based on the calculated Q value and the target Q value, the parameters θ of the Q evaluation network are updated using the stochastic gradient descent algorithm. The update formula is:

[0079]

[0080] Subsequently, the central controller will assign parameter θ t+1 This is sent to each user device to update the local Q-evaluation network. Additionally, the parameters θ of the Q-target network in the central controller are also updated. t - The Q-evaluation network is updated once every N time steps during training, based on the parameters θ of the Q-evaluation network. t You can get it by copying directly.

[0081] This application has achieved the following beneficial effects:

[0082] 1. This invention adopts an architecture that combines centralized training with distributed execution. During the training phase, global network state information is used to perform joint policy optimization to improve the overall performance of the system. During the execution phase, each agent makes decisions autonomously based on local observation results, reducing dependence on the central control unit. This invention reduces signaling transmission overhead and decision latency, and enhances the real-time response capability and scalability of the system.

[0083] 2. By introducing a gated loop unit into the network structure of the switching agent and utilizing its memory mechanism in the time dimension to effectively integrate historical observation information, this invention enhances the agent's ability to perceive the evolution trend of the environmental state and improves the accuracy and stability of switching decisions in dynamic scenarios.

[0084] 3. A novel exploration mechanism is employed, which introduces trainable parameterized noise into the connection weights of the neural network. This noise intensity adaptively adjusts during training, generating exploration perturbations relevant to the current policy during action selection. This mechanism avoids the waste of samples caused by completely random exploration, achieving directional and efficient exploration. It improves sample utilization while accelerating policy convergence, making it particularly suitable for network switching decision optimization in high-dimensional action spaces and dynamic environments.

[0085] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description

[0086] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0087] Figure 1This is a flowchart illustrating a method for switching end-to-end network slices in a heterogeneous cellular network according to an embodiment of the present invention.

[0088] Figure 2 This is a schematic diagram of a heterogeneous cellular network architecture based on end-to-end network slicing provided in an embodiment of the present invention;

[0089] Figure 3 This is a schematic diagram of the framework of the MA-DDGQN algorithm provided in an embodiment of the present invention;

[0090] Figure 4 This is a schematic diagram of the handover process based on MA-DDGQN in a heterogeneous cellular network scenario provided by an embodiment of the present invention. Detailed Implementation

[0091] To make the objectives, technical solutions, and advantages of the present invention clearer, preferred embodiments of the present invention will be described below in conjunction with the accompanying drawings, providing a further detailed explanation of the invention. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention.

[0092] This invention provides a distributed network slicing switching method based on a multi-agent deep recurrent dual-Q network. For example... Figure 1 The diagram illustrates a distributed network slicing switching method based on a multi-agent deep recurrent dual-Q network. The specific implementation includes:

[0093] S1: Construct a heterogeneous cellular network scenario based on end-to-end network slicing;

[0094] S2: Define the service types of network slices and establish the service model and switching model required for network slice switching issues;

[0095] S3: The handover problem is modeled as a distributed partially observable Markov decision process model. In the heterogeneous cellular network based on end-to-end network slicing, the slice handover intelligently selects the appropriate network slice according to the actual needs of the terminal device and the network conditions to meet the different network service needs of the terminal device.

[0096] S4: The distributed Markov decision process model is optimized and solved using a multi-agent deep recurrent double-Q network algorithm to obtain a distributed handover strategy. Then, based on the distributed handover strategy, slice distributed handover is performed on the heterogeneous cellular network system based on end-to-end network slices.

[0097] The heterogeneous cellular network scenario based on end-to-end network slicing constructed in this invention is as follows: Figure 2 As shown, the system comprises M macro base stations (MBS), P microcell base stations (PBS), and F millicell base stations (FBS), with each base station represented as B = {b1, ..., b}.k ,…,b K In this scenario, N network slices are deployed, and the slices are represented as N = {n1, ..., n}. i ,…,n N U terminal devices are distributed within the area and move randomly. The terminal devices are represented as U = {u1, ..., u}. j ,…,u U Each network slice can cover M′ physical base stations, and a single physical base station can deploy multiple network slices. All NSs share the same physical transmission resources, including the RAN's radio transmission bandwidth and power, as well as the transmission bandwidth in the core network. Therefore, there are a total of M′×N base station-slice pairs in this model, represented as H={h1,…,h l ,…,h M′×N}, where l=(i-1)×M′+k′ is the id of the base station-slice pair, representing slice n i The k′-th base station is covered. The radio transmission resources of each base station will be allocated to all network slices deployed on it as needed.

[0098] The MA-DDGQN algorithm framework diagram is as follows: Figure 3 As shown, based on the MA-DDGQN algorithm, a dual deep Q-network is used as the local decision agent for each terminal device in the implementation of the switching strategy. According to the distributed partially observable Markov decision process model established above, within the management scope of the same centralized controller, all terminal devices have the same action set, local observation information, and reward function, thus allowing for a unified agent structure design. Based on this uniformity, the system centrally stores the interaction data of all local agents in an experience replay buffer, from which the centralized controller samples training samples for centralized training of the DDQN network. Simultaneously, each terminal device has a local Q-network with the same structure as the Q-network trained in the centralized controller, used for autonomous policy reasoning and action selection. Overall, the MA-DDGQN algorithm adopts an architecture of centralized training and distributed execution.

[0099] To illustrate the steps of the algorithm in detail, the MA-DDGQN algorithm is described below using one time step in the training process as an example.

[0100] 1) Distributed experience collection phase: At time step t, terminal device u j Obtain the current local state observation. j,t Choose an action a using the ε-greedy strategy as follows. j,t All terminal devices select actions and obtain a combined action a. t =(a 1,t ,...,aj,t ,...,a U,t Then the terminal device performs the selected action, namely, switching the base station-slice pair, and receives the corresponding reward r. j,t After the combined action is executed, the state of the entire system changes from s t Convert to s t+1 Meanwhile, terminal device u j Obtain new local state observations. j,t+1 In this time step, each user device will obtain an experience tuple d. i =(o j,t ,a j,t ,r j,t ,o j,t+1 The training samples are stored in the experience playback buffer D in the central controller.

[0101] 2) Centralized DDGQN training phase: First, the centralized controller randomly selects a small batch from the experience replay buffer D. It contains B experiences and d x =(o x ,a x ,r x ,o x′ Then, the central controller will... x As input to the neural network, the hidden state H is obtained after processing by the GRU layer. t Finally, the hidden state H will be... t After being processed by the Dropout layer, it is used as input to the fully connected layer to calculate the Q-value and Q-target value for each sample.

[0102] 3) Parameter update stage: Based on the calculated Q value and Q target value, the parameters θ of the Q evaluation network are updated using the stochastic gradient descent algorithm. t Subsequently, the central controller will assign parameter θ t+1 This is sent to each user device to update the local Q-evaluation network. Additionally, the parameters θ of the Q-target network in the central controller are also updated. t - The Q-evaluation network is updated once every N time steps during training, based on the parameters θ of the Q-evaluation network. t You can get it by copying directly.

[0103] During the testing phase, each terminal device autonomously selects its actions based on the pre-trained Q-network and an e-greedy policy. In actual deployment, the MA-DDGQN algorithm can be pre-trained and directly applied to real communication systems. Because each terminal device relies solely on its local observations for decision-making during testing, without transmitting any local data to the central controller, the MA-DDGQN algorithm effectively protects user privacy.

[0104] In this embodiment of the invention, the handover process of the MA-DDGQN handover mechanism is further illustrated. Its specific implementation process is as follows: Figure 4 As shown, this process is achieved through the collaborative work of user equipment, base stations, and software-defined network controllers.

[0105] At the start of each time step, the centralized controller broadcasts the current system status to all terminal devices. Subsequently, each terminal device, based on the received system status and its own status information, independently makes a handover decision through its respective evaluation network and sends the generated handover request back to the centralized controller. After collecting requests from all UEs, the centralized controller forwards the handover request to the corresponding target base station and target SDN controller. Once the target and source SDN controllers receive and confirm the request, they jointly complete the handover operation for that UE. After the handover is complete, the source SDN controller releases the resources previously occupied by the base station and network slices. After the handover process for all UEs is completed, the centralized controller updates the system status.

[0106] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented, in whole or in part, as a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0107] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A distributed network slice handover method based on multi-agent deep recurrent double Q network, characterized in that, The method comprises the following steps: S1: constructing an end-to-end network slice-based heterogeneous cellular network scenario; S2: defining the service type of the network slice and establishing the service model and switching model required for the network slice switching problem; S3: modeling the switching problem as a distributed partially observable Markov decision process model, and intelligently selecting a suitable network slice in the end-to-end network slice-based heterogeneous cellular network according to the actual demand of the terminal device and the network condition to meet the different network service demands of the terminal device; S4: obtaining a distributed switching strategy by optimizing and solving the distributed Markov decision process model based on a multi-agent deep recurrent double Q network algorithm, and then performing slice distributed switching on the end-to-end network slice-based heterogeneous cellular network system based on the distributed switching strategy.

2. The distributed network slice handover method based on multi-agent deep recurrent double Q network according to claim 1, characterized in that: In S1, the end-to-end network slice-based heterogeneous cellular network scenario comprises a radio access network, a core network, and a bearer network. The radio access network comprises terminal devices, macro base stations, small base stations, home base stations, and network slices. Different types of network slices are deployed on one or more base stations, one base station is deployed with one or more types of network slices, and terminal devices are connected to network slices on various types of base stations through the radio access network to obtain related network services. The core network mainly comprises network function virtualization (NFV) and software-defined network (SDN). NFV reduces network cost and improves network flexibility, and SDN realizes dynamic configuration and management of network functions. The resource state of the bearer network and the core network is abstracted as the available bandwidth of the core network. In S1, the end-to-end network slice has been deployed in the initialization stage of the network according to the SLA predefined by different services, and the association relationship between the base station and the network slice remains unchanged during the simulation process. In terms of user demand, two parameters are used to describe the service quality requirements of the terminal device, including: represents a network slice n i a minimum threshold of guaranteed transmission rate; T i , indicates the network slice n i The maximum time that can guarantee the transmission rate below the minimum threshold value; In order to serve as many terminal devices as possible, it is assumed that each network slice provides service to terminal devices with the minimum transmission rate required by the service.

3. The distributed network slice handover method based on multi-agent deep recurrent double Q network according to claim 1, characterized in that: In S2, the service type of the network slice is defined, and the service model and switching model required for the network slice switching problem are established. The task model specifically includes: The task model adopts a full buffer service mode, i.e., each terminal device (UE) in the network always has a data stream to be transmitted. In order to ensure that the service quality provided by the network slice remains consistent, two parameters are considered to describe the service quality requirements of the terminal device in terms of user demand, including: a minimum threshold value indicative of a transmission rate; τ j , represents the bearable time, i.e. the terminal device u j allows its transmission rate to be lower than the minimum threshold for a maximum time; To serve as many terminal devices as possible, it is assumed that each network slice serves a terminal device with the minimum transmission rate required by the service, so at time step t the terminal device u j The access BS-NS pair h l The estimated radio transmission network bandwidth consumed is calculated as follows: Furthermore, the transmission rate of the terminal device u j through the base station b k access BS-NS pair h l at time t and is calculated as: where W B denotes the bandwidth obtained by each user, is the signal-to-interference-plus-noise ratio between the terminal device u j and the base station b k at time step t, and, in addition, is calculated as: where p i represents the transmit power of the base station b i , G i,j (t) represents the radio channel gain between the terminal device u j and the base station b k , I represents the interference of other base stations to the terminal device u j , and N0 represents the Gaussian white noise. For convenience of description, at time step t, respectively use and to represent the core network bandwidth and the radio access network bandwidth required to ensure the user service quality requirement of the task of the terminal device u j when accessing the BS-NS pair h l , wherein 4. The distributed network slice handover method based on multi-agent deep recurrent double Q network according to claim 1, characterized in that: In S2, the service type of the network slice is defined, and the service model and switching model required for the network slice switching problem are established. The switching model specifically includes: The switching model: the slice switching action is divided into three cases: different base stations same slice switching, different base stations different slice switching, and same base station different slice switching. The first two cases are considered as different base station switching, and the last case is considered as switching within the same base station. In order to facilitate the expression of the problem, it is assumed that all terminal devices make switching decisions at each time point. However, if the base station-network slice pair connected at the previous time point is selected, the actual switching process will not start, because the terminal device moves randomly and the channel quality deteriorates. In this case, the terminal device u j does not consume the resources of the BS-NS pair h l , and the delay will increase by 1 when the terminal device u j continues to delay, and: terminal device u j An interruption occurs at time step t, where γ j (t0) is the terminal device u j The real-time transmission rate obtained, in addition, if the terminal device triggers a handover procedure at the current time step, it will be moved to the end of the UE queue.

5. The distributed network slice handover method based on multi-agent deep recurrent double Q network according to claim 1, characterized in that: In S3, the switching problem is modeled as a distributed partially observable Markov decision process model. In the end-to-end network slice-based heterogeneous cellular network, the slice switching intelligently selects a suitable network slice according to the actual demand of the terminal device and the network condition to meet the different network service demands of the terminal device, specifically including: The handover problem is modeled as a distributed partially observable Markov decision process model, specifically including: modeling the network slice handover problem as a distributed partially observable Markov decision process model to maximize the average cumulative reward related to the profit of serving terminal devices, handover cost and interruption penalty as the optimization target, and the distributed partially observable Markov decision process model can be represented by a tuple (I, S, A, P, R, O, γ) to represent its mathematical model, and the detailed definitions of specific states, actions and rewards are as follows: The global state includes both the network state and the local state; therefore, at time step t, the global state is defined as s. G,t =[s N,t ,s U,t ]∈S, where s N,t and s U,t They are defined as follows: s U,t = [s 1,t ,...,s j,t ,...,s U,t ] where s j,t is the local state of terminal k j , defined as: where I j is the terminal u j At the time of decision, the id of the base station-slice pair accessed in the current time step, d j The state vector of terminal u j The available time remaining before the service interruption occurs, since the association between base stations and network slices (BS-NS) is configured to remain static for a certain period of time, the state vector dimensions of all user equipment in this area are completely consistent, based on this, the local observation of terminal u j is defined o j,t = [s N,t ; s j,t ] At each decision time t, the terminal device needs to select one from the available base station-slice combinations for access, and the action is defined as the selection behavior of the terminal device to the base station-slice pair, so A={h l} l=1,2…K′×S , where h l is the id of the base station-slice pair, and each terminal device needs to select one at time step t to access the BS-NS pair; terminal u j The action a at time step t is defined as: a j,t =h l ; Due to the stability of the base station-slice association relationship, the selectable action set of all terminals in the region is completely the same, therefore, the joint action at time t can be expressed as the tuple of all terminal actions: a t = (a 1,t ,…,a j,t ,…,a U,t ) To define the reward function of terminal u j , three types of utility are designed: the terminal device served utility R1, the switching cost R2, and the interruption penalty R3. Given the difference in the decisions of each terminal device in the switching process, the reward of terminal device u j under the current decision can be expressed as: Thus, the optimization goal of the handover decision process is the cumulative reward obtained by all terminal devices in each time step, which is expressed as 6. The distributed network slice handover method based on multi-agent deep recurrent double Q-network according to claim 1 and claim 5, characterized in that: In the S4, the distributed Markov decision process model is optimized and solved by using a multi-agent deep recurrent double Q network algorithm to obtain a distributed handover strategy, and then the slice distributed handover is performed on the end-to-end network slice based heterogeneous cellular network system based on the distributed handover strategy, wherein the parameterized noise exploration specifically includes: The parameterized noise is added to the full connection layer of the neural network, and the noise parameter ε w and ε b is added to the connection weight of the full connection layer, so that the full connection layer not only needs to learn the mean μ of the weight w, but also needs to learn its variance σ, and both the mean and the variance are learned as network parameters. After introducing noise in the full connection layer, the output Q function will be randomized, thereby improving the exploration ability of the agent. Decomposed Gaussian noise is used to generate noise ε. It is assumed that the full connection layer generates p independent Gaussian noises ε i according to the number of neurons, and the next layer generates q independent Gaussian noises ε j according to the number of neurons. Therefore, p+q noises will be finally generated in these two layers, and each and can be represented as: where f is a function used in the present application 7. The distributed network slice handover method based on multi-agent deep recurrent double Q-network according to claim 1 and claim 5, characterized in that: In the S4, the distributed Markov decision process model is optimized and solved by using a multi-agent deep recurrent double Q network algorithm to obtain a distributed handover strategy, and then the slice distributed handover is performed on the end-to-end network slice based heterogeneous cellular network system based on the distributed handover strategy, wherein the recurrent neural network specifically includes: GRU maintains the same input-output structure as traditional RNNs, with the input s at each time point being... t The state h of the previous moment t-1 They work together in the current calculation process, where h t-1 Carrying information from previous points in time, GRU controls the flow of information by introducing update and reset gates, enabling the network to retain or discard information at different time steps. Its structure is as follows: t Represents the update gate, which is used to control the hidden state h of the previous time step. t-1 To what extent should it be preserved, and what is the current candidate hidden state h? t To what extent should it affect the hidden state h at the current moment? t The calculation method for the update gate is defined as follows: z t = σ(W sz s t + W hz h t-1 + b z ) where σ represents a Sigmoid activation function, W represents weights, r t represents a reset gate, determined by the current input s t and the previous hidden state h t-1 , which functions to adjust the degree of influence of past hidden state information on the current state, and if the current input is irrelevant to previous information, r t can play a role to remove the influence of h t-1 , and the reset gate is calculated as: r t = σ(W sr s t + W hr h t-1 ) hidden state h t is z t , h t-1 and h t ' are calculated, when z t is close to 1, h t-1 is approximately equal to h t , the new hidden state almost completely retains the h t-1 information of the previous moment, reflecting a high degree of dependence on past information, on the contrary, when z t is close to 0, the new state is mainly determined by the current input s t , reducing the influence of the past hidden state h t-1 , this flexible information processing mechanism makes GRU both capture long-term dependencies and adapt to short-term changes, the calculation method of hidden state is defined as: h t = (1 - z)h t + z t ⊙ h t-1 candidate hidden state h t r by resetting the gate t and the previous hidden state h t-1 element-wise multiplication, and the current input s t h′ after being processed by a tanh function, the range of h t ′ is between -1 and 1, when approaching 0, the result of element-wise multiplication is close to 0, which means that the information of the previous state is almost completely ignored, so that the GRU can focus on the current input and adapt to capture short-term dependencies, on the contrary, when r t approaches 1, h t ′ retains a large amount of information of the past state h t-1 , which is beneficial to maintain long-term memory, and the calculation method of the candidate hidden state is defined as: h t ′ = Tanh(W s s t + W h (r t ⊙ h t-1 ) + b n ) In the MA-DDGQN algorithm model, a GRU neural network is added before the full connection layer of the Q value network on the basis of the DQN algorithm, and then the Q value is predicted according to the historical information h t and the current state s t , h t is the input returned by the network at the previous time step, and then the MA-DDGQN algorithm uses the Q value network to calculate the Q value of the action and obtains the optimal action corresponding to the maximum Q value, and puts the optimal action into the target value network for Q value calculation, that is, first find the optimal action by using the Q value network, and the corresponding formula is as follows: where h t - represents the hidden state output by the GRU neural network after processing the next state, as the input of the full connection layer, θ t represents the Q value network parameters, Q(h t - , θ) represents a series of Q values of the next state, and the meaning of this formula is to obtain the optimal action corresponding to the next state, and then use the selected action a max (h t - , θ t ) to calculate the target Q value in the target network, that is: y = r + γQ'(h t - a max (h - ,θ t ) ; θ t - ) where y represents the target Q value, r represents the reward of the environment feedback, γ is a discount factor, θ t - is the parameter of the target value network, and after the calculated Q value and the Q target value, the parameter θ of the Q evaluation network is updated by using a stochastic gradient descent algorithm, and the update formula is: Subsequently, the centralized controller sends the parameters θ t+1 to each user device to update the local Q evaluation network, and in addition, the parameters θ t - The Q evaluation network is trained for one update at N time steps, and the parameters θ t are copied directly.

8. The distributed network slice handover method based on multi-agent deep recurrent double Q network according to claim 1, characterized in that, The handover process of the MA-DDGQN handover mechanism includes: through the cooperative work of the terminal device, the base station and the software defined network controller, at the beginning of each time step, the centralized controller broadcasts the current system state to all terminal devices, then each terminal device independently makes handover decisions through its own evaluation network according to the received system state and its own state information, and sends the generated handover request back to the centralized controller, after collecting all terminal device requests, the centralized controller will pass the handover request to the corresponding target base station and target software defined network controller, when the target and source software defined network controllers receive the request and confirm, they will jointly complete the handover operation of the terminal device, after the handover is completed, the source software defined network controller will release the resources occupied by the original base station and network slice, after all terminal device handover processes are executed, the centralized controller will update the system state. 9.A computer readable storage medium, comprising instructions which, when executed on a computer, cause the computer to perform a distributed network slice handover method based on multi-agent deep recurrent double Q network according to claim 1. 10.A control system for implementing the distributed network slice handover method based on multi-agent deep recurrent double Q network according to claim 1.