Resource configuration and optimization method of 5G network slice communication system
By applying deep reinforcement learning and adaptive hierarchical duel DQN algorithm in 5G network slices, the problem that resource allocation is difficult to meet heterogeneous business needs is solved, and efficient resource utilization and user satisfaction are maximized.
Patent Information
- Application Number
- CN202510273661.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-17
AI Technical Summary
When 5G network slicing faces multiple heterogeneous business needs, it is difficult to achieve efficient allocation and optimization of resources, resulting in difficult to ensure user satisfaction and service quality.
Deep reinforcement learning methods, especially the adaptive hierarchical duel DQN algorithm, are used to optimize the resource configuration of 5G network slices. This method dynamically allocates users to the base station by building a system model of a hierarchical architecture, and allocates resources according to service quality requirements, geographical location and real-time load conditions.
It achieves maximum user satisfaction under the limitations of restricted resources and slicing service quality, ensures efficient resource utilization and rapid adaptability to dynamic network conditions, and improves the flexibility and customization capabilities of 5G network slicing.
Smart Images

Figure CN120166423A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for resource allocation and optimization in a communication system, and more specifically to a method for resource allocation and optimization in a 5G network slice communication system. Background Art
[0002] The rapid development of mobile communication technology, especially the deployment of 5G, has driven the transformation from traditional physical infrastructure to a flexible and programmable virtualization paradigm. Network Function Virtualization (NFV) and Software Defined Network (SDN) have become key enablers for addressing the requirements of manageability, flexibility, and automation. NFV virtualizes hardware functions on commercial off-the-shelf platforms into software-based virtual network functions, enhancing scalability, while SDN decouples the control plane and data plane for centralized orchestration. However, NFV faces latency challenges in sensitive applications, and SDN struggles to achieve quality-of-service-aware traffic management under high loads.
[0003] Network slicing is a 5G innovation enabled by the integration of SDN and NFV, representing a transformative paradigm. SDN manages global control and resource orchestration, while NFV manages the lifecycle of virtual network functions. This synergy allows for the dynamic provisioning of services and allocation of resources based on slice-specific service level agreements, thus supporting scenario-optimized service customization. By enabling multiple logically isolated virtual networks on a shared infrastructure, resource utilization and service differentiation are enhanced, providing end-to-end connectivity for various applications.
[0004] The network slicing architecture is currently divided into three main types: Ultra-Reliable Low-Latency Communication (URLLC) slices, Enhanced Mobile Broadband (eMBB) slices, and Voice over Long-Term Evolution (VoLTE) slices. URLLC slices are designed for mission-critical applications that require exceptional reliability (>99.999%) and ultra-low latency (<1ms), such as industrial control systems, remote surgery, autonomous vehicle networks, and smart grid management. eMBB slices are optimized for high-throughput data transmission scenarios, supporting high-definition video streaming, immersive media experiences (VR / AR), 4K / 8K broadcast systems, and bandwidth-intensive enterprise solutions. VoLTE slices are specifically used to provide high-definition voice services, while ensuring isolation from data traffic interference during voice communication bandwidth allocation. This hierarchical slicing paradigm enables granular resource partitioning based on heterogeneous service requirements, thereby providing highly flexible and customizable network services. This feature is particularly valuable in a multi-tenant environment where different service providers can coexist on a shared infrastructure while maintaining strict performance isolation and quality-of-service guarantees. Summary of the Invention
[0005] The purpose of the present invention is to propose a resource allocation and optimization method for a 5G network slice communication system. This method applies the deep reinforcement learning method to the resource allocation of network slices, faces the 5G network, promotes the transformation from traditional physical infrastructure to a flexible and programmable virtualization mode, thereby enhancing the scalability of network resources, and proposes a hierarchical slice mode to support fine resource partitioning based on heterogeneous service requirements, so as to provide highly flexible and customized network services.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A resource allocation and optimization method for a 5G network slice communication system, comprising the following steps:
[0008] S1: Build a hierarchical architecture 5G network slice system model and a time slot communication model. The framework of the present invention is implemented in an orthogonal frequency division multiple access downlink network controlled by software-defined network (SDN) including multiple base stations. Among them, a centralized downlink transmission framework is proposed, integrating multiple base stations, network slices and users under the control of an SDN controller. In the 5G network slice, it is usually divided into three categories. According to different categories, they provide different services for users, and the required quality of service requirements are also different. For resource allocation, it is centrally controlled by an SDN controller. In a communication cycle, each user generates a service demand with a certain fixed probability, and each service demand corresponds to a slice. The SDN controller will make the following configurations according to the demands generated by users and their geographical locations: select a base station for each user to communicate with; allocate network resources to each base station, and then further allocate them to the users connected to it through the base station. This method realizes a hierarchical resource allocation mechanism, in which the controller dynamically allocates users to base stations according to slice-specific quality of service requirements, geographical proximity and real-time load conditions, including the number of base stations, the type and number of network slices, the number of users requesting services, and communication model-related parameters, etc.;
[0009] S2: In the system model constructed according to S1, analyze the obtained data transmission rate and throughput. Under relevant constraint conditions, according to the 3GPP standard, a frame can be divided into multiple time slots, and users generate service demands at the beginning of each frame. The SDN allocates resources for each user in each time slot. Since 5G uses orthogonal frequency division multiple access modulation technology, the resources in this patent refer to time-frequency resource blocks (RBs). Our measurement criteria are delay and throughput. Define the satisfaction of delay and the satisfaction of throughput respectively, and the satisfaction of users is defined as the weighted sum of the two. Since each slice has different requirements for delay and throughput, the weights we allocate are also different. Finally, we establish an optimization problem of maximizing the total user satisfaction under limited resources and slice quality of service constraints;
[0010] S3: The optimization problem is essentially non-convex, and there is a strong correlation between each time slot. These characteristics cannot be solved by traditional convex optimization methods. Therefore, the formulated optimization problem is transformed into an equivalent Markov decision process, and the state space, action space, and reward function adapted to the system environment are set. The Dueling DQN algorithm in the classical reinforcement learning algorithm is used to seek the dynamic solution. Given the complex state and action spaces involved, we propose an Adaptive Hierarchical Dueling DQN algorithm. This algorithm systematically decomposes the optimization task into sequential sub-problems, and this architecture supports efficient bandwidth allocation; at the same time, the priority of the user satisfaction index is determined through reward shaping and hierarchical policy learning. There are mainly three innovations: maximizing user satisfaction by coordinating resource scheduling; using multi-dimensional state representation to capture the spatial user distribution and slice type diversity, so as to realize the intelligent partitioning of cross-base station network resources; ensuring efficient resource utilization while maintaining rapid adaptability to dynamic network conditions.
[0011] Furthermore, the hierarchical architecture 5G network slice system model established in S1 includes the following steps:
[0012] S11: Our framework is implemented in an SDN-controlled orthogonal frequency division multiple access downlink network containing multiple base stations. Users are evenly distributed in a circular coverage area with a radius of D. In each service period, each user probabilistically generates a service request for a specific network slice. The SDN controller executes a logically ordered resource allocation process, including two consecutive stages. The allocation mechanism first selects the optimal base station for the user according to the user's geographical location and network slice requirements; subsequently, the controller allocates spectrum resources to each user through the allocated base station, and the allocation parameters are clearly controlled by the spatial distribution characteristics and slice-specific quality of service requirements;
[0013] S12: We assume represents the set of base stations, represents the set of network slices. For each slice the associated user group is defined as where M n represents the number of users requesting slice n service;
[0014] S13: Following the 3GPP specification, our 5G network architecture adopts a time-slot communication model, where each time slot with a duration of T is divided into K mini-slots, where the duration of each mini-slot is T min = T / K, and the total system bandwidth F is evenly divided into C sub-channels, and the bandwidth of each sub-channel is According to the principle of orthogonal frequency division multiple access, this configuration generates KC orthogonal resource blocks (RBs) in the time-frequency grid. The SDN controller allocates spectrum resources to the base stations. The minimum bandwidth granularity per time slot is 1 RB. Let represent the total spectrum resources allocated to base station y in the k-th time slot;
[0015] S14: In the second stage, the SDN controller performs fine-grained resource allocation at the user layer, where represents the bandwidth allocated to user m n and served by base station y during time slot k. The achievable transmission rate of user m n in the k-th time slot can be expressed as:
[0016]
[0017] where represents the transmit power allocated by the base station to user m in the k-th time slot n ; represents the distance between user m n and base station y. N0 represents the noise power spectral density. The squared channel gain follows an exponential distribution with unit parameter, and δ represents the path loss exponent of the sub-channel;
[0018] S15: For user in network slice n, the effective throughput is constrained by the specific slice delay requirement K n . When k > K n , its spectrum resources are unavailable. Therefore, the cumulative effective throughput of the system in time slot k can be expressed as:
[0019]
[0020] Furthermore, the optimization problem of establishing the corresponding hierarchical structure according to the system model described in S2 includes the following steps:
[0021] S21: We establish a dual quality-of-service detection metric, considering both system delay and throughput requirement parameters. The throughput satisfaction of user m n can be expressed as:
[0022]
[0023] where R n represents the minimum throughput threshold of slice n;
[0024] S22: The delay satisfaction of user m n can be expressed as:
[0025]
[0026] wherein represents the amount of data generated by user m n ;
[0027] S23: The main challenge of this method is to allocate resources to simultaneously meet the latency requirements and throughput requirements of all slices, so as to maximize the user satisfaction across the system. To solve the heterogeneous quality-of-service priority problem, we introduce slice-specific weighted coefficients α n (throughput) and β n (latency), so the formulated global satisfaction metric can be expressed as:
[0028] This configuration allows the primary quality-of-service requirements of each user to be preferentially satisfied during resource allocation, thus achieving the optimization of the global quality of service of the entire system;
[0029] S24: Based on the above analysis, we establish the following constrained optimization problem, which can be expressed as maximizing the overall user satisfaction by dynamically allocating bandwidth resources across time slots within a service period:
[0030]
[0031] Furthermore, the constraint conditions of the optimization problem described in S24 are specifically explained as follows: the aggregated bandwidth allocated to users under each base station shall not exceed the bandwidth allocated to this base station by the SDN for each time period; the cumulative bandwidth allocated to all users in the entire system is still constrained by the total available bandwidth F; the SDN enforces the sub-channel level regime when allocating bandwidth to users, where the bandwidth is allocated in integer multiples of the sub-channel unit.
[0032] Furthermore, converting the optimization problem described in S3 into a Markov decision process for solution includes the following steps:
[0033] S31: List the state space in the Markov decision process of the optimization problem;
[0034] S32: List the action space in the Markov decision process of the optimization problem, which consists of two consecutive resource management decision processes by the agent within time slot k;
[0035] S33: The immediate reward within time slot k is calculated as the total satisfaction of all users, weighted by their respective priorities and quality-of-service requirements, and its reward function can be expressed as:
[0036]
[0037] S34: Our goal is to maximize the long-term cumulative reward, which can be expressed as:
[0038]
[0039] Among them, π is the policy for managing resource allocation decisions, T is the optimization scope consistent with the slice lifecycle constraints, and γ is the discount factor.
[0040] Furthermore, listing the state space in the Markov decision process of the optimization problem described in S31 includes the following steps:
[0041] S311: The communication cycle is divided into k time slots: k;
[0042] S312: The slice deadline of each slice network: K n ;
[0043] S313: The number of users of each slice: M n ;
[0044] S314: The distance between user m n and base station y:
[0045] S315: The task workload generated by user m n :
[0046] S316: The effective throughput of user m n in each time slot k:
[0047] Therefore, at time slot t, the state s(t) can be formally expressed as:
[0048]
[0049] Furthermore, listing the action space in the Markov decision process of the optimization problem described in S32 includes the following steps:
[0050] S321: Base station allocation:
[0051]
[0052] Among them, represents the serial number index of the base station assigned to user m n ;
[0053] S322: Bandwidth allocation:
[0054]
[0055] Among them, represents the amount of resources assigned to user m n ;
[0056] S323: Therefore, at time slot k, the state a(t) can be formally expressed as:
[0057]
[0058] Furthermore, a resource allocation and optimization method for a 5G network slice communication system. The algorithm described in S35 of S3 is called an adaptive hierarchical dueling algorithm execution process, including the following steps:
[0059] S35: An execution process of an adaptive hierarchical dueling algorithm, characterized by including the following steps:
[0060] S351: The algorithm starts to execute, input and initialize system environment area parameters including user equipment, number of base stations, communication cycle, small time slots, channel bandwidth, and number of slices, etc.;
[0061] S352: Initialize deep reinforcement learning network parameters, including greedy policy parameter ∈, experience replay buffer capacity, and target network soft update parameters, etc.;
[0062] S353: Reset the environmental variables in the network, and put the SDN samples into the minimum experience buffer, observe the current state space and input it into the value network;
[0063] S354: The system allocates base station and bandwidth resources to users;
[0064] S355: Judge t ≤ K n , if the condition holds, calculate the user throughput and return to execute S353; otherwise execute S356;
[0065] S356: Obtain the action space, and calculate the reward value (including user satisfaction and effective throughput);
[0066] S357: Obtain the state tuple at the next moment, and store it in the experience replay pool;
[0067] S358: Complete the gradient descent operation steps, update the value network parameters, and perform a soft update on the target network;
[0068] S359: Judge If the condition holds, directly break the internal loop and end the algorithm process to S3510; otherwise execute S353;
[0069] S3510: Output bandwidth: When the SDN allocates bandwidth for users, it allocates in integer multiples of the sub-channel unit, thereby ending the algorithm process.
[0070] The present invention has the following advantages and effects:
[0071] 1. The present invention proposes a centralized downlink transmission framework that integrates multiple base stations, network slices, and users under the control of an SDN controller. The framework implements a hierarchical resource allocation mechanism, where the controller dynamically assigns users to base stations according to slice-specific quality of service requirements, geographical proximity, and real-time load conditions. Resource blocks are then allocated to the base stations and subsequently assigned to slice-specific users.
[0072] 2. The present invention formulates a non-convex optimization problem to maximize user satisfaction by coordinating resource scheduling and develops an adaptive hierarchical dueling algorithm to address the optimization challenges. The algorithm uses a multi-dimensional state representation to capture the spatial user distribution and slice-type diversity, thereby enabling intelligent partitioning of network resources across base stations. This design ensures efficient resource utilization while maintaining rapid adaptability to dynamic network conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 is a flowchart of the implementation process of the method of the present invention;
[0074] Figure 2 is a system model diagram constructed by the method of the present invention;
[0075] Figure 3 is a flowchart of the algorithm execution. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0076] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0077] The following describes the specific implementation of the present invention in detail in conjunction with specific embodiments.
[0078] As Figure 1 and Figure 2 shown, a resource configuration and optimization method for a 5G network slice communication system according to an embodiment of the present invention includes the following steps:
[0079] S1: Build a hierarchical architecture 5G network slicing system model and a time-slot communication model. The framework of the present invention is implemented in an orthogonal frequency division multiple access downlink network controlled by software-defined network (SDN) that includes multiple base stations. A centralized downlink transmission framework is proposed, integrating multiple base stations, network slices, and users under the control of one SDN controller. In 5G network slicing, it is usually divided into three categories. According to different categories, they provide different services to users, and the required quality of service requirements vary. For resource allocation, it is centrally controlled by one SDN controller. Within one communication cycle, each user generates a service demand with a certain fixed probability, and each service demand corresponds to one slice. The SDN controller will make the following configurations according to the demands generated by users and their geographical locations: select a base station for each user to communicate with; allocate network resources to each base station, and then further allocate them to the users connected to the base station through the base station. This method realizes a hierarchical resource allocation mechanism, where the controller dynamically allocates users to base stations according to slice-specific quality of service requirements, geographical proximity, and real-time load conditions, including the number of base stations, the type and number of network slices, the number of users requesting services, and communication model-related parameters, etc.;
[0080] Specifically, in the above S1, the hierarchical architecture 5G network slicing system model is as follows:
[0081] S11: Our framework is implemented in an orthogonal frequency division multiple access downlink network controlled by SDN that includes multiple base stations. Users are evenly distributed in a circular coverage area with a radius of D. Within each service cycle, each user probabilistically generates a service request for a specific network slice. The SDN controller executes a logically ordered resource allocation process, including two consecutive stages. The allocation mechanism first selects the optimal base station for users according to their geographical locations and network slice demands; subsequently, the controller allocates spectrum resources to each user through the allocated base station, and the allocation parameters are clearly controlled by spatial distribution characteristics and slice-specific quality of service requirements;
[0082] S12: We assume represents the set of base stations, represents the set of network slices. For each slice the associated user group is defined as where M n represents the number of users requesting slice n service;
[0083] S13: Following the 3GPP specification, our 5G network architecture adopts a time-slot communication model, where each time slot with a duration of T is divided into K mini-slots, where the duration of each mini-slot is T min = T / K, and the total system bandwidth F is evenly divided into C sub-channels, and the bandwidth of each sub-channel is According to the principle of orthogonal frequency division multiple access, this configuration generates KC orthogonal resource blocks (RBs) in the time-frequency grid. The SDN controller allocates spectrum resources to the base stations. The minimum bandwidth granularity for each time slot is 1 RB. Let represent the total spectrum resources allocated to base station y in the k-th time slot;
[0084] S14: In the second stage, the SDN controller performs fine-grained resource allocation at the user layer, where represents the bandwidth allocated to user m n and served by base station y during time slot k. User m n can achieve a transmission rate within the k-th time slot, which can be expressed as:
[0085]
[0086] where represents the transmit power allocated by the base station to user m within the k-th time slot n , represents the distance between user m n and base station y. N0 represents the noise power spectral density, and the squared channel gain follows an exponential distribution with a unit parameter, and δ represents the path loss exponent of the sub-channel;
[0087] S15: For user in network slice n, the effective throughput is constrained by the specific slice delay requirement K n . When k > K n , its spectrum resources are unavailable. Therefore, the cumulative effective throughput of the system within time slot k can be expressed as:
[0088]
[0089] S2: In the system model constructed according to S1, the data transmission rate and throughput are analyzed. Under relevant constraints, according to the 3GPP standard, a frame can be divided into multiple time slots, and user service demands are generated at the beginning of each frame. The SDN allocates resources to each user for each time slot. Since 5G adopts orthogonal frequency division multiple access modulation technology, the resources in this patent refer to time-frequency resource blocks (RBs). Our measurement criteria are delay and throughput. The satisfaction of delay and throughput are defined respectively, and the user satisfaction is defined as the weighted sum of the two. Since each slice has different requirements for delay and throughput, the weights we allocate are also different. Finally, we establish an optimization problem to maximize the total user satisfaction under limited resources and slice quality of service constraints;
[0090] Specifically, in step S2, an optimization problem is established based on the system model, and the specific optimization problem is as follows:
[0091] S21: We establish a dual quality-of-service detection index, considering two parameters of system latency and throughput requirements simultaneously. The throughput satisfaction of user m n can be expressed as:
[0092]
[0093] where R n represents the minimum throughput threshold of slice n;
[0094] S22: The latency satisfaction of user m n can be expressed as:
[0095]
[0096] where represents the amount of data generated by user m n ;
[0097] S23: The main challenge of this method is to allocate resources to simultaneously meet the latency requirements and throughput requirements of all slices, thereby maximizing the user satisfaction within the system. To solve the heterogeneous quality-of-service priority problem, we introduce slice-specific weighting coefficients α n (throughput) and β n (latency). Therefore, the formulated global satisfaction index can be expressed as:
[0098] This configuration allows the primary quality-of-service requirements of each user to be preferentially satisfied during resource allocation, thereby achieving the optimization of the global quality of service of the entire system;
[0099] S24: Based on the above analysis, we establish the following constrained optimization problem, which can be expressed as maximizing the overall user satisfaction by dynamically allocating bandwidth resources across time slots within the service period:
[0100]
[0101] The constraint conditions of the optimization problem described in S24 are specifically explained as follows: The aggregated bandwidth allocated to users under each base station shall not exceed the bandwidth allocated to the base station by the SDN for each time period; The cumulative bandwidth allocated to all users in the entire system is still constrained by the total available bandwidth F; The SDN enforces the sub-channel level system when allocating bandwidth to users, where the bandwidth is allocated in integer multiples of sub-channel units.
[0102] S3: The optimization problem constitutes an essentially non-convex problem, and there is a strong correlation between each time slot. These characteristics cannot be solved by traditional convex optimization methods. Therefore, the formulated optimization problem is transformed into an equivalent Markov decision process, and a state space, an action space, and a reward function that adapt to the system environment are set. The Dueling DQN algorithm in the classical reinforcement learning algorithm is used to seek a dynamic solution. Given the complex state and action spaces involved, we propose an adaptive hierarchical dueling DQN algorithm. This algorithm systematically decomposes the optimization task into sequential sub-problems, and this architecture supports efficient bandwidth allocation; at the same time, the priority of the user satisfaction index is determined through reward shaping and hierarchical policy learning. There are mainly three innovations: maximizing user satisfaction by coordinating resource scheduling; using multi-dimensional state representation to capture the spatial user distribution and slice type diversity, so as to realize the intelligent partitioning of cross-base station network resources; ensuring efficient resource utilization while maintaining fast adaptability to dynamic network conditions.
[0103] Specifically, in S3, according to the conversion of the problem into a Markov process, its state space is listed as follows:
[0104] S31: List the state space in the Markov decision process of the optimization problem. The specific state space includes the following:
[0105] S311: The communication cycle is divided into k time slots: k;
[0106] S312: The slice deadline of each slice network: K n ;
[0107] S313: The number of users of each slice: M n ;
[0108] S314: User m n The distance between and base station y:
[0109] S315: The task workload generated by user m n Generated:
[0110] S316: The effective throughput of user m within each time slot k n Effective throughput:
[0111] Therefore, at time slot t, the state s(t) can be formally expressed as:
[0112]
[0113] Specifically, in step S3, the action space in the optimization problem of the Markov decision process is listed. The agent consists of two consecutive resource management decision-making processes within time slot k, including the following components:
[0114] S32: List the state space in the optimization problem of the Markov decision process. The specific action space includes the following:
[0115] S321: Base station allocation:
[0116]
[0117] Among them, represents the serial number index of the base station allocated to user m n ;
[0118] S322: Bandwidth allocation:
[0119]
[0120] Among them, represents the amount of resources allocated to user m n ;
[0121] S323: Therefore, at time slot k, the state a(t) can be formally expressed as:
[0122]
[0123] Specifically, in step S3, according to the problem being converted into a Markov process, its reward function is listed as follows:
[0124] S33: The immediate reward within time slot k is calculated as the total satisfaction of all users, weighted according to their respective priorities and quality of service requirements. Its reward function can be expressed as:
[0125]
[0126] S34: Our goal is to maximize the long-term cumulative reward, which can be expressed as:
[0127]
[0128] where π is the policy for managing resource allocation decisions, T is the optimization scope consistent with the slice lifecycle constraints, and γ is the discount factor.
[0129] The above is only the preferred embodiment of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can also be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent.
[0130] Specifically, as Figure 3 shown, the execution process of an adaptive hierarchical dueling algorithm proposed in S35 of S3 is as follows:
[0131] S351: The algorithm starts to execute, and inputs and initializes system environment area parameters including user equipment, number of base stations, communication cycle, mini-slot, channel bandwidth, and number of slices, etc.;
[0132] S352: Initialize the deep reinforcement learning network parameters, including greedy policy parameter ∈, experience replay buffer capacity, and target network soft update parameter, etc.;
[0133] S353: Reset the environmental variables in the network, and put the SDN samples into the minimum experience buffer, observe the current state space and input it into the value network;
[0134] S354: The system allocates base station and bandwidth resources to users;
[0135] S355: Judge t ≤ K n , if the condition holds, calculate the user throughput and return to execute S353; otherwise execute S356;
[0136] S356: Obtain the action space, and calculate the reward value (including user satisfaction and effective throughput);
[0137] S357: Obtain the state tuple at the next moment, and store it in the experience replay pool;
[0138] S358: Complete the gradient descent operation steps, update the value network parameters, and perform a soft update on the target network;
[0139] S359: Judge If the condition holds, directly break the inner loop to end the algorithm process to S3510; otherwise execute S353;
[0140] S3510: Output bandwidth: When the SDN allocates bandwidth for users, it allocates in integer multiples of the sub-channel unit, thereby ending the algorithm process.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention. Without departing from the concept of the present invention, several deformations and improvements can also be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent.
Claims
1. A resource configuration and optimization method for a 5G network slicing communication system, characterized in that: The method comprises the following steps: S1: Build a network slicing system model and time slot communication model for layered 5G. This framework is implemented in an orthogonal frequency division multiple access downlink network controlled by SDN (software defined network) containing multiple base stations. A centralized downlink transmission framework is proposed, which integrates multiple base stations, network slices and users under the control of an SDN controller. In 5G network slices, they are usually divided into three categories. According to different categories, they provide different services to users and have different service quality requirements. For resource allocation, it is centrally controlled by an SDN controller. In a communication cycle, each user generates service requirements with a fixed probability, and each service requirement corresponds to a slice. The SDN controller will perform the following configurations based on the needs and geographical location generated by the users: select a base station for each user to communicate with; allocate network resources to each base station, and then further allocate them to users connected to it through the base station; this method implements a layered resource allocation mechanism, in which the controller dynamically allocates users to base stations based on slice-specific service quality requirements, geographical proximity and real-time load conditions, including the number of base stations, the type and number of network slices, the number of users requesting services and communication model related parameters; S2: According to the system model constructed in S1, the data transmission rate and throughput are analyzed. Under relevant constraints, according to the 3GPP standard, a frame can be divided into multiple time slots, and users generate service requirements at the beginning of each frame; SDN allocates resources to each user in each time slot. Since 5G adopts orthogonal frequency division multiple access modulation technology, this resource refers to the time-frequency resource block (RB); the measurement criteria are delay and throughput, and the satisfaction of delay and throughput are defined respectively. The user satisfaction is defined as the weighted sum of the two; since each slice has different requirements for delay and throughput, the weights allocated are also different. Finally, an optimization problem of maximizing total user satisfaction under limited resources and slice service quality constraints is established; S3: The optimization problem constitutes an essentially non-convex problem, and there is a strong correlation between each time slot. These characteristics cannot be solved by traditional convex optimization methods, so the formulated optimization problem is transformed into an equivalent Markov decision process, and the state space, action space and reward function that adapt to the system environment are set, and the DuelingDQN algorithm in the classical reinforcement learning algorithm is used to seek a dynamic solution; in view of the complex state and action space involved, an adaptive hierarchical dueling DQN algorithm is proposed; the algorithm systematically decomposes the optimization task into sequential sub-problems, and this architecture supports efficient bandwidth allocation; at the same time, the priority of user satisfaction indicators is determined through reward shaping and hierarchical strategy learning; there are three points: maximize user satisfaction by coordinating resource scheduling; use multi-dimensional state representation to capture spatial user distribution and slice type diversity, thereby realizing intelligent partitioning of network resources across base stations; ensure efficient resource utilization while maintaining rapid adaptability to dynamic network conditions.
2. According to claim 1, a resource configuration and optimization method for a 5G network slicing communication system is characterized in that: The network slicing system model for building a layered 5G architecture described in S1 includes the following steps: S11: The framework is implemented in an SDN-controlled OFDMA downlink network containing multiple base stations, where users are evenly distributed in a circular coverage area of radius D. In each service period, each user probabilistically generates a service request for a specific network slice; The SDN controller performs a logically ordered resource allocation process consisting of two consecutive phases. The allocation mechanism first selects the optimal base station for the user based on the user's geographical location and network slice requirements. Subsequently, the controller allocates spectrum resources to individual users through the assigned base stations, with allocation parameters explicitly controlled by spatial distribution characteristics and slice-specific quality of service requirements. S12: Assumptions represents the set of base stations, Represents a collection of network slices; for each slice The associated user group is defined as Among them, M n Indicates the number of users requesting services from slice n.
3. According to claim 1, a resource configuration and optimization method for a 5G network slicing communication system is characterized in that: The parameters related to the time slot communication model described by S1 include the following steps: S13: Following the 3GPP specification, the 5G network architecture adopts a time slot communication model, where each time slot of duration T is divided into K small time slots, where The duration of each mini-slot is T min =T / K, the total system bandwidth F is evenly divided into C sub-channels, and the bandwidth of each sub-channel is According to the principle of orthogonal frequency division multiple access, this configuration generates KC orthogonal resource blocks (RBs) in the time-frequency grid. The SDN controller allocates spectrum resources to base stations. The minimum bandwidth granularity of each small time slot is 1 RB. represents the total spectrum resources allocated to base station y in the kth mini-slot; S14: In the second stage, the SDN controller performs fine-grained resource allocation at the user level, where Indicates the value assigned to user m n bandwidth, is served by base station y during time slot k, and user m n In the kth time slot, the achieved transmission rate can be expressed as: in Indicates that the base station allocates to user m in the kth time slot n Transmit power, Indicates user m n The distance from the base station y, N0 represents the noise power spectral density, and the square channel gain It follows an exponential distribution with unit parameter, and δ represents the path loss exponent of the subchannel; S15: For users In a network slice n, the effective throughput is subject to the specific slice latency requirement K n Constraints; when k>K n When , its spectrum resources are not available, so the cumulative effective throughput of the system in time slot k is expressed as:
4. The resource configuration and optimization method of a 5G network slicing communication system according to claim 1, characterized in that: S2 describes that under relevant constraints, with the goal of maximizing the satisfaction of all users in the system, a corresponding optimization problem is established, including the following steps: S21: Dual service quality detection indicators are established, taking into account both system delay and throughput requirements. n The throughput satisfaction can be expressed as: Among them, R n Indicates the minimum throughput threshold of slice n; S22: User m n The delayed satisfaction can be expressed as: in Indicates user m n The amount of data generated; S23: The challenge of this method is to allocate resources to meet the latency requirements and throughput requirements of all slices at the same time, so as to maximize system-wide user satisfaction; in order to address the heterogeneous service quality priority problem, a slice-specific weighting coefficient α is introduced n (throughput) and β n (delay), so the global satisfaction index is expressed as: This configuration allows to prioritize each user’s primary QoS requirements during resource allocation, thus optimizing the global QoS of the entire system; S24: Based on the above analysis, the following constrained optimization problem is established to maximize the overall user satisfaction by dynamically allocating bandwidth resources across time slots within the service period, expressed as: The optimization problem is subject to three constraints: the aggregate bandwidth allocated to users under each base station must not exceed the bandwidth allocated to the base station by SDN for each time period; the cumulative bandwidth allocated to all users in the entire system is still constrained by the total available bandwidth F; SDN enforces a sub-channel level system when allocating bandwidth to users, in which bandwidth is allocated in integer multiples of sub-channel units.
5. According to claim 1, a resource configuration and optimization method for a 5G network slicing communication system is characterized in that: S3 transforms the optimization problem into a Markov decision process for solution, including the following steps: S31: List the state space of the Markov decision process for the optimization problem, including the following components: S311: The communication cycle is divided into k time slots: k; S312: Slice duration of each slice network: K n ; S313: Number of users per slice: M n ; S314: User m n Distance from base station y: S315: User m n The resulting task workload: S316: In each time slot k, user m n Effective throughput: S317: Therefore, at time slot t, the state s(t) is formally expressed as: S32: The action space in the Markov decision process of the optimization problem is listed, which consists of two consecutive resource management decision processes of the agent in time slot k, including the following components: S321: Base station allocation: in, Indicates the value assigned to user m n The serial number index of the base station; S322: Bandwidth allocation: in, Indicates the value assigned to user m n The amount of resources; S323: Therefore, at time slot k, the state a(t) is formally expressed as: S33: The instant reward in time slot k is calculated as the total satisfaction of all users, weighted by their respective priorities and service quality requirements, and its reward function is expressed as: S34: The goal is to maximize the long-term cumulative reward, expressed as: where π is the policy that manages resource allocation decisions, T is the optimization horizon consistent with the slice lifecycle constraints, and γ is the discount factor.
6. According to claim 1, a resource configuration and optimization method for a 5G network slicing communication system is characterized in that: The algorithm is called an adaptive layered duel algorithm execution process, which includes the following steps: S35: An adaptive layered duel algorithm execution process, comprising the following steps: S351: The algorithm starts to execute, inputs and initializes system environment area parameters including user equipment, number of base stations, communication cycle, mini-slot, channel bandwidth and number of slices; S352: Initialize the deep reinforcement learning network parameters, including the greedy strategy parameter ∈, the experience revisit buffer capacity and the target network soft update parameter; S353: reset the environment variables in the network, put the SDN sample into the minimum experience buffer, observe the current state space and input the value network; S354: The system allocates base stations and bandwidth resources to the user; S355: Determine t≤K n If the condition is met, the user throughput is calculated and the process returns to S353; otherwise, S356 is executed; S356: Obtain action space and calculate reward value (including user satisfaction and effective throughput); S357: Obtain the state tuple of the next moment and store it in the experience replay pool; S358: Complete the gradient descent operation step, update the value network parameters, and perform soft update on the target network; S359: Judgment If the condition is met, the internal loop is directly broken and the algorithm process ends to S3510; Otherwise, execute S353; S3510: Output bandwidth: When allocating bandwidth to users, SDN allocates it in integer multiples of sub-channel units, thereby ending the algorithm process.