A resource scheduling method for hybrid business throughput optimization

By constructing a throughput-weighted model and multi-agent reinforcement learning, the resource scheduling strategy was optimized, solving the resource scheduling problem in the mixed service scenario of URLLC and mMTC, and improving network transmission efficiency and data volume.

CN115633402BActive Publication Date: 2025-10-28CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211302266.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2025-10-28
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

In foggy cells, under mixed URLLC and mMTC service scenarios, different QoS requirements make it difficult to schedule resources reasonably, resulting in a decrease in the amount of data transmitted over the network.

Method used

A throughput-weighted model is constructed, which is combined with multi-agent reinforcement learning. Through iterative search using deep reinforcement learning, resource scheduling strategies are optimized to solve resource allocation problems in mixed business scenarios.

Benefits of technology

It improved network throughput performance and resource utilization, met the transmission needs of different services, and enhanced the overall transmission efficiency of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115633402B_ABST
    Figure CN115633402B_ABST
Patent Text Reader

Abstract

This invention relates to a resource scheduling method for optimizing throughput in mixed services, belonging to the field of communication technology. The method includes: S1: constructing a channel model for a mixed service transmission system based on the data characteristics of URLLC and mMTC services; S2: constructing a throughput optimization model for mixed services in a foggy cell; S3: solving the mixed service throughput optimization model using a multi-agent approach, i.e., iteratively finding the optimal resource scheduling strategy for throughput under mixed services using a multi-agent reinforcement learning model. This invention can effectively improve network throughput performance and has broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology and relates to a resource scheduling method for optimizing throughput of mixed services. Background Technology

[0002] With the continuous development of 5G mobile communication technology, intelligent lifestyles and production methods have become a trend. Data generated by diverse devices is transmitted across the network to where it is needed, with data generated by IoT devices accounting for a significant portion, and smart buildings and industrial automation being the main growth drivers. IoT services in these scenarios are characterized by small individual data packets but large overall traffic volumes, posing new challenges to current networks. To meet the increasing communication demands of IoT services, current wireless networks need continuous evolution.

[0003] In traditional wireless access networks, data processing is primarily handled by the deployed base stations (BS), necessitating the deployment of numerous BS devices to meet communication demands. This not only increases construction costs but also results in low spectrum utilization. Therefore, to improve spectrum efficiency and achieve higher network performance, China Mobile has introduced cloud computing technology into its wireless access networks, proposing the Cloud Radio Access Networks (C-RAN) architecture. In the C-RAN architecture, multiple base band units (BBUs) are aggregated in the cloud to form a BBU pool, and specific virtualization technologies enable flexible allocation of centralized resources. Simultaneously, remote radio heads (RRHs) are deployed close to the user side to ensure regional signal coverage, and their operating status can be dynamically adjusted based on RRH load, thereby improving network performance.

[0004] To further improve the quality of localized services in wireless networks, Cisco proposed the concept of fog networks based on fog computing principles. Researchers then combined fog networks with wireless networks to form the F-RAN (fog radio access networks) architecture. In the F-RAN architecture, fog access nodes (F-APs) can be deployed at the network edge to form a large number of service nodes that can provide communication, computing, and storage capabilities, thereby distributing the information processing pressure from the network to the edge. Because F-APs in F-RAN can cache some data, they have significant advantages in alleviating fronthaul link load and improving network performance. Furthermore, since F-APs have fog computing capabilities, they can perform radio signal processing and resource management, thus having a natural advantage over C-RAN and H-CRAN architectures in improving the service efficiency of local services and increasing resource utilization.

[0005] In the 5G era, the International Telecommunication Union (ITU) classifies application scenarios into three categories: Massive Machine-Type Communications (mMTC), Ultra Reliable Low Latency Communications (URLLC), and Enhanced Mobile Broadband (eMBB). Among these, mMTC is the primary application scenario for IoT services, encompassing areas such as smart buildings and smart cities. IoT applications in this scenario need to meet the connectivity demands of massive numbers of IoT devices, transmit more IoT services, and assist in decision-making by collecting data sensed by IoT devices to improve decision effectiveness. However, with the increasing number of access network devices, different devices require different communication resources, posing significant challenges to network resource scheduling. Therefore, resource scheduling tailored to the characteristics of different services is crucial for improving network performance. Summary of the Invention

[0006] In view of this, the purpose of this invention is to provide a resource scheduling method for optimizing throughput in mixed service scenarios. This method addresses the problem of reduced network data transmission volume caused by the difficulty in rationally scheduling resources due to different QoS requirements in mixed URLLC and mMTC service scenarios in foggy cells. Based on the characteristic of URLLC services generating small data packets, a weighted modeling analysis of throughput in mixed service scenarios is performed. Then, a joint sub-channel allocation and power control method is designed to improve the data transmission volume of URLLC small data packet services and mMTC services. Simultaneously, the throughput weighting model and the sub-channel allocation and power control method are constructed into a multi-agent reinforcement learning problem. Deep reinforcement learning is used to handle the resource scheduling problem under different channel conditions in mixed service scenarios, seeking the optimal resource allocation decision.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A resource scheduling method for optimizing throughput for mixed services is proposed. This method addresses the communication needs of IoT-based mixed services in F-RAN by comprehensively considering the data characteristics, QoS requirements, and channel conditions of different services to optimize network throughput performance. First, the data characteristics of URLLC and mMTC services are analyzed, along with the system channel conditions when these two services are transmitted in combination. Second, to optimize network throughput performance, a throughput weighting model for mixed services is constructed. Finally, to find the optimal power allocation method and channel selection in unknown environments, the throughput optimization problem for mixed services is transformed into an optimal policy solution problem in unknown environments using multi-agent reinforcement learning, based on the throughput weighting model. The iterative search method of deep reinforcement learning is then used to find the optimal resource scheduling strategy for different services in different environments.

[0009] The method specifically includes the following steps:

[0010] S1: Construct a channel model for a hybrid service transmission system based on the data characteristics of URLLC and mMTC services;

[0011] S2: Construct a throughput optimization model for mixed services in fog cells;

[0012] S3: Solve the mixed service throughput optimization model using multi-agent methods, that is, use a multi-agent reinforcement learning model to iteratively find the resource scheduling strategy with the optimal throughput under mixed services.

[0013] Furthermore, step S1 specifically includes the following steps:

[0014] S11: In a fog cell, all IoT tasks generated by IoT devices are collected by the scheduler. The F-AP transmits the channel information collected in the previous cycle to the scheduler. The scheduler combines this channel information to allocate the current IoT task transmission queue and make resource allocation decisions for each IoT task.

[0015] S12: In a fog cell, the fog access node F-AP transmits H URLLC services, and there are J mMTC services within the coverage area of ​​the F-AP, which are communication between IoT devices. Among them, the URLLC services obtain higher transmission rate services by connecting with the F-AP, while the mMTC services are mainly information exchange between some IoT devices. Since IoT devices are mainly single-antenna devices, it is assumed that all devices use a single antenna in the scenario. Then, the set of URLLC services in the scenario can be represented as H = {0,…,h}, and the set of mMTC services can be represented as J = {0,…,j}.

[0016] S13: In fog cells, downlink transmission is the primary consideration. Therefore, OFDM technology is used as the background, and it is assumed that the fading of the channel is roughly the same within a sub-bandwidth, and that different sub-bandwidths are independent of each other. In a time slot interval, if URLLC service and mMTC service share the transmission sub-channel, then URLLC service will suffer from interference from mMTC service, that is, interference from IoT device transmitter to URLLC service. Similarly, mMTC service will also suffer from interference from URLLC service.

[0017] Traditional channel capacity calculations based on Shannon's theorem assume an ideal capacity with infinite code length and a bit error rate (BER) approaching zero. However, in real-world environments, the BER ε is non-zero, and the code length transmitted through the channel is finite. Literature has shown through theoretical and experimental analysis that longer code lengths can effectively increase transmission rates. Furthermore, adding checksums and other verification information can ensure reliable transmission and achieve a lower BER. However, IoT services require the transmission of numerous small data packets. If each IoT service requires adding significant redundant information to ensure high-reliability transmission, it would result in the transmission of a large amount of useless information in the network, thus reducing information utility. Therefore, considering latency and information effectiveness, URLLC services are analyzed based on small data packet characteristics, while mMTC services are analyzed based on Shannon capacity.

[0018] URLLC service throughput C in the network URLLC Represented as:

[0019]

[0020] in, B0 is the code length occupied by the URLLC service, T is the transmission duration, C is the channel capacity per unit length in infinite code length (i.e., the Shannon limit), V is the channel dispersion, and Q is the channel length of the URLLC service. -1 (·) is a function The inverse function of , where ε is the bit error rate;

[0021] The throughput C of mMTC services in the network mMTC Represented as:

[0022]

[0023] in, Code length occupied by mMTC services.

[0024] Furthermore, step S2 specifically includes the following steps:

[0025] S21: While meeting the requirements of low latency and high reliability in URLLC communication links, it can also increase the overall system capacity, meaning that more mMTC service information can be transmitted. However, according to S1, when mixed services are transmitted in a coordinated manner, they will interfere with each other, so it is necessary to construct an mMTC service admission control model.

[0026]

[0027] Among them, C j,t C represents the communication link capacity of the j-th mMTC service within time slot t; mMTC,t The amount of traffic that the mMTC service needs to transmit within time slot t, where n is the sub-bandwidth sequence number, N is the total number of sub-bandwidths, and ρ j,h When the h-th URLLC service and the j-th mMTC service are in the same sub-bandwidth, the bandwidth allocation coefficient for the mMTC service.

[0028] S22: Combining the data characteristics of different services and jointly considering the transmission power control of different services, a throughput optimization model for mixed services in fog cells is constructed:

[0029]

[0030]

[0031]

[0032]

[0033]

[0034] Where, n max P represents the maximum number of symbols that can be used within a time slot. h,maxP represents the maximum transmission power for all URLLC services within the total bandwidth. j,max P represents the maximum transmission power of all mMTC services within the total bandwidth. h,n The power required to transmit the h-th URLLC service within the n-th sub-bandwidth; P j,n The power required to transmit the j-th mMTC service within the n-th sub-bandwidth; ρ j,n Assign control coefficients to sub-bandwidths, ρ j,n ∈{0,1}, when ρ j,n When ρ = 1, it means that the j-th mMTC service can use the n-th sub-bandwidth for transmission. j,n =0 indicates that the nth sub-bandwidth cannot be used for transmission; ρ j',n Assign control coefficients to the sub-bandwidths of the remaining mMTC services, j'∈J, and j'≠j; if ρ j',n =1 indicates that more mMTC services can be added to the current sub-bandwidth for transmission.

[0035] Furthermore, step S3 specifically includes: In a hybrid service coexistence system, mMTC services and URLLC services can be transmitted together within the same sub-bandwidth. Because different services require different transmission power and latency, the resource scheduling problem for different services is modeled as a multi-agent reinforcement learning problem. Reinforcement learning can iteratively find the resource scheduling strategy with optimal throughput under different environments, satisfying the transmission needs of as many services as possible, thereby increasing the service capacity in the network. Each communication link between URLLC and mMTC services is considered an agent, collecting experience from different decisions in different environments to select and execute a suitable scheduling strategy in the current environment. When multiple services need to be transmitted simultaneously, multiple communication links form a multi-agent cluster, jointly exploring sub-bandwidth scheduling and power control strategies in the current environment.

[0036] Furthermore, step S3 specifically includes the following steps:

[0037] S31: Constructing a multi-agent reinforcement learning model:

[0038] Intelligent agent: Each communication link in the network;

[0039] State: The main network performance parameters considered in this scenario include sub-channel operating state, remaining communication link capacity, service tolerance latency, and channel conditions within the sub-bandwidth; therefore, the network state can be defined as s = {s state ,s capacity ,s time ,s SNR}∈S, where s state For sub-bandwidth occupancy status, s capacityFor the remaining bandwidth capacity, s time Let c be the service tolerance delay, c be the channel signal-to-noise ratio, and S be the state space set.

[0040] Action: For the resource scheduling problem in a hybrid service coexistence system, each agent can decide to use any sub-bandwidth for transmission, and the transmission power within that sub-bandwidth; considering the downlink transmission power of F-AP, the transmission power control for mMTC services is defined as P. mMTC ={50,100,150,200}mW; therefore, the action space has a dimension of 4*H, with each action corresponding to a sub-bandwidth and transmission power selection; although the action space of a single communication link is not very large, the overall action space will become large when the number of communication links is large. Therefore, the transmission power is not set as a continuous variable.

[0041] S32: Solve for the optimal policy in the multi-intelligence reinforcement learning model, and implement the resource scheduling policy with optimal throughput under mixed business by designing a suitable reward function and error function;

[0042] Reward Function: To obtain an optimized resource scheduling strategy in the network, a reward function needs to be set. Each agent can adjust its scheduling strategy based on the reward value obtained from each decision, thereby gradually approaching the optimal decision. The main objective is to improve network throughput performance and mMTC service transmission success rate in mixed service scenarios within a time interval T. The reward function for the mMTC service communication link is set as follows:

[0043]

[0044] Among them, C max For Shannon capacity, C re-mMTC λ represents the remaining mMTC traffic volume; μ1 is the URLLC communication link agent factor, indicating whether there is a URLLC service to be transmitted within the current sub-bandwidth; μ2 is the mMTC communication link agent factor, indicating whether there is an mMTC service to be transmitted within the current sub-bandwidth; λ is a hyperparameter used to adjust the network's willingness to access mMTC services.

[0045] The reward function for the entire model is set as follows:

[0046]

[0047] Where, ω URLLC ω is a weighting factor used to adjust the transmission volume of URLLC services. mMTC This is a weighting factor used to adjust the transmission volume of mMTC services;

[0048] The error function is defined as follows:

[0049]

[0050] Where γ is the loss factor, s t Let a be the state at time t. t Let A be the action at time t, and let A be the set of actions, θ be the action at time t. t Let θ' be the parameter before updating in the stochastic gradient descent method. t Q represents the updated parameters in the stochastic gradient descent method. t The empirical value at time t.

[0051] The beneficial effects of this invention are as follows: Addressing the problem of low resource scheduling efficiency in hybrid service cooperative transmission within foggy cells, this invention provides a resource scheduling strategy for optimizing hybrid service throughput. By analyzing the data characteristics of URLLC and mMTC services, and the channel model for hybrid cooperative transmission, a throughput-weighted model for hybrid services is constructed. Furthermore, a multi-agent reinforcement learning model is used to solve for the optimal resource scheduling strategy for hybrid service cooperative transmission. The method provided by this invention can effectively improve network throughput performance and has broad application prospects.

[0052] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0053] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0054] Figure 1 This is a diagram of the multi-layer collaborative resource scheduling framework of the present invention;

[0055] Figure 2 This is a diagram of the multi-agent resource scheduling architecture designed for this invention. Detailed Implementation

[0056] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0057] Please see Figures 1-2 , Figure 1 The diagram illustrates a resource scheduling method for optimizing throughput in mixed-service environments, which includes the following steps:

[0058] Step 1: Analysis of URLLC and mMTC Service Data Characteristics and Construction of System Channel Model: Since URLLC and mMTC services are mixed in foggy cells, and URLLC includes small data packet services transmitting security information, while simultaneously meeting ultra-low latency and ultra-high reliability requirements, performance analysis of different services under limited code length is necessary. The preferred approach includes the following steps:

[0059] Step 1.1: In a fog cell, all IoT tasks generated by IoT devices are collected by the scheduler. The F-AP transmits the channel information collected in the previous cycle to the scheduler. The scheduler combines this channel information to allocate the current IoT task transmission queue and make resource allocation decisions for each IoT task.

[0060] Step 1.2: Define a fog cell where the F-AP transmits H URLLC services and J mMTC services exist within the F-AP's coverage area, representing communication between IoT devices. URLLC services achieve higher transmission rates by connecting to the F-AP, while mMTC services primarily involve information exchange between IoT devices. Since IoT devices mainly use single antennas, this scenario assumes all devices use a single antenna. Therefore, the set of URLLC services in the scenario can be represented as H = {0,…,h}, and the set of mMTC services can be represented as J = {0,…,j}.

[0061] Step 1.3: In foggy cells, downlink transmission is the primary consideration. Therefore, OFDM technology is used as the background, and it is assumed that the fading of the channel is roughly the same within a sub-bandwidth, and that different sub-bandwidths are independent of each other. In a time slot interval, if URLLC service and mMTC service share the transmission sub-channel, then URLLC service will suffer from SINR interference from mMTC service, that is, interference from IoT device transmitter to URLLC service. Similarly, mMTC service will also be affected by interference from URLLC service.

[0062]

[0063] Among them, P h,n The power required to transmit the h-th URLLC service within the n-th sub-bandwidth; P j,n The power required to transmit the j-th mMTC service within the n-th sub-bandwidth; g h For channel coefficients; g j,nρ represents the interference coefficient caused to the channel by transmitting the j-th mMTC service within the n-th sub-bandwidth; j,n Assign control coefficients to sub-bandwidths, ρ j,n ∈{0,1}, when ρ j,n When ρ = 1, it means that the j-th mMTC service can use the n-th sub-bandwidth for transmission. j,n =0 indicates that the nth sub-bandwidth cannot be used for transmission; σ 2 This represents the noise power spectral density.

[0064]

[0065] Where, ρ j',n Assign control coefficients to the sub-bandwidth of the remaining mMTC services; if ρ j',n =1 indicates that more mMTC services can be added to the current sub-bandwidth for transmission.

[0066] URLLC services are small data packet services. Therefore, a theoretical analysis of the channel transmission rate under finite code length is performed, and the channel rate (bits / s / Hz) under finite code length is obtained:

[0067]

[0068] Where n is the code length; C is the unit channel capacity in infinite code length, i.e., the Shannon limit; V is the channel dispersion V = 1 - (1 + γ). -2 Where γ is the signal-to-noise ratio; Q -1 (·) is a function The inverse function of .

[0069] The throughput of URLLC services in the network can then be expressed as:

[0070]

[0071] Where B0 is the size of the acquired spectrum; T is the transmission duration;

[0072] The throughput of mMTC services in the network can be expressed as:

[0073]

[0074] Step 2: Construction of a Weighted Throughput Model for Mixed Services: Based on the analysis of transmission performance and channel conditions for different services in the mixed service scenario in Step 1, a weighted throughput model for mixed services was designed to improve network throughput performance and meet the access needs of more services. The preferred model includes the following steps:

[0075] Step 2.1: While satisfying the requirements of low latency and high reliability transmission in the URLLC communication link, increasing the overall system capacity means that more mMTC service information can be transmitted. However, as mentioned in Step 1, mixed service collaborative transmission can cause mutual interference, so an mMTC service admission control model needs to be constructed:

[0076]

[0077] Among them, C j,t C represents the communication link capacity of the j-th mMTC service within time slot t; mMTC,t The amount of traffic that mMTC services need to transmit within time slot t;

[0078] Step 2.2: Combining the data characteristics of different services in Step 1, and jointly considering the transmission power control of different services, construct a throughput optimization model for mixed services in a fog cell:

[0079]

[0080]

[0081]

[0082]

[0083]

[0084] Where, n max The maximum number of symbols that can be used within a time slot; P h,max P represents the maximum transmission power for all URLLC services within the total bandwidth. j,max This represents the maximum transmission power for all mMTC services within the total bandwidth.

[0085] Step 3: Multi-Agent Hybrid Service Throughput Optimization Strategy: In a hybrid service coexistence system, mMTC and URLLC services can be transmitted simultaneously within the same sub-bandwidth. Because different services require different transmission power and latency, the resource scheduling problem for different services can be modeled as a multi-agent reinforcement learning problem. Reinforcement learning can iteratively find better resource scheduling strategies under different environments to meet the transmission needs of more services as much as possible, thereby improving the service capacity of the network. Each communication link between URLLC and mMTC services is considered an agent, which collects experience from different decisions in different environments and selects a suitable scheduling strategy to execute in the current environment. When multiple services need to be transmitted simultaneously, multiple communication links form a multi-agent cluster to jointly explore sub-bandwidth scheduling and power control strategies in the current environment. The optimal strategy includes the following steps:

[0086] Step 3.1: In a multi-agent reinforcement learning model, such as Figure 2 As shown, the decision-making process of each agent can be defined as a set (S, A, P, r, d), where S is the set of state spaces; A is the set of action spaces; and P is the state transition probability, which refers to the probability of transitioning from the current state s to the current state d. t If action a is taken... t At that time, a new state s is obtained. t+1 The probability of success is given by r; the reward value is given by d; since all agents receive the same reward, as it is necessary to simultaneously complete URLLC service transmissions and access as many mMTC services as possible; d is the loss factor. During model learning, each agent can obtain the corresponding reward value after taking an action, and then update the reinforcement learning model and gradually find better decisions. When the model training is complete, each agent, after obtaining the state information of the current environment, selects an action that brings greater benefits based on the historical experience obtained from the trained model.

[0087] Step 3.3: Based on the analysis in Step 3.1 and combined with Step 2, construct a multi-agent model:

[0088] Intelligent agent: Each communication link in the network.

[0089] State: The main network performance parameters considered in this scenario include sub-channel operating state, remaining communication link capacity, service tolerance latency, and channel conditions within the sub-bandwidth. Therefore, the network state can be defined as s = {s state ,s capacity ,s time ,s SNR}∈S,s state Sub-bandwidth occupancy status; s capacity For remaining bandwidth capacity; stime c is the service tolerance latency; c is the channel signal-to-noise ratio.

[0090] Action: For the resource scheduling problem in a hybrid service coexistence system, each agent can decide to use any sub-bandwidth for transmission, and the transmission power within that sub-bandwidth. Considering the downlink transmission power of F-AP, the transmission power control for mMTC services is defined as P. mMTC ={50,100,150,200}mW. Therefore, the action space has a dimension of 4*H, with each action corresponding to a sub-bandwidth and transmission power selection. Although the action space of a single communication link is not very large, the overall action space will become large when the number of communication links is large. Therefore, the transmission power is not set as a continuous variable.

[0091] Reward Function: To obtain an optimized resource scheduling strategy in the network, a reward function needs to be set. Each agent can adjust its scheduling strategy based on the reward value obtained from each decision, thereby gradually approaching the optimal decision. The main objective is to improve network throughput performance and mMTC service transmission success rate in mixed service scenarios within a time interval T. The reward function for the mMTC service communication link is set as follows:

[0092]

[0093] Among them, C max λ is the Shannon capacity; μ1 is the URLLC communication link agent factor, indicating whether there is a URLLC service to be transmitted within the current sub-bandwidth; μ2 is the mMTC communication link agent factor, indicating whether there is an mMTC service to be transmitted within the current sub-bandwidth; λ is a hyperparameter used to adjust the network's willingness to access mMTC services.

[0094] The reward function for the entire model is set as follows:

[0095]

[0096] Where, ω URLLC ω is a weighting coefficient used to adjust the transmission volume of URLLC services; mMTC This is a weighting coefficient used to adjust the transmission volume of mMTC services.

[0097] Step 3.3: In reinforcement learning, each agent chooses a policy π to maximize the cumulative reward. Here, policy π refers to the probability distribution of the agent's current state s that maps to action a. The cumulative discount function is typically used to represent the expected reward at policy π, and is defined as follows:

[0098]

[0099] Where, ξt This refers to the loss rate;

[0100] The goal in the learning process is to find optimization strategies. The definition is as follows:

[0101]

[0102] When found When, it means that the current state s can be obtained. t The optimization strategy is as follows:

[0103]

[0104] In order to find optimization strategies Iterative algorithms can be used to find the state transition probability P(s). However, in real-world environments, due to a lack of prior knowledge, it is difficult to obtain the state transition probability P(s). t+1 |s t ,a t Therefore, we utilize the Deep Q-Network (DQN) model from reinforcement learning to address the lack of experience in unknown environments.

[0105] Each agent has a DQN, taking the state space S as input and outputting the value function corresponding to all actions A. The DQN model is trained through multiple iterations. In each iteration, all agents employ a soft policy: selecting the action with the maximum estimated value in the state-action space with probability 1-ε, and selecting a random action with probability ε. When the channel state changes or the environment changes due to agent actions, each agent collects and stores the current state-action space, reward value, and the state space for the next step in an experience pool L. In each iteration, some information is extracted from the experience pool to update the parameters θ in the stochastic gradient descent method. Through multiple iterations, a fixed set of parameters is obtained to reduce the error value. When the loss factor is γ, the error function is defined as follows:

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A resource scheduling method for optimizing throughput in mixed services, characterized in that, The method specifically includes the following steps: S1: Construct a channel model for a hybrid service transmission system based on the data characteristics of URLLC and mMTC services; S2: Construct a throughput optimization model for mixed services in fog cells; S3: Solve the hybrid service throughput optimization model using multi-agent methods, that is, use a multi-agent reinforcement learning model to iteratively find the resource scheduling strategy with the optimal throughput under hybrid services, specifically including the following steps: S31: Constructing a multi-agent reinforcement learning model: Intelligent agent: Each communication link in the network; State: The network state is defined as s = {s state ,s capacity ,s time ,s SNR }∈S, where s state For sub-bandwidth occupancy status, s capacity For the remaining bandwidth capacity, s time Let c be the service tolerance delay, c be the channel signal-to-noise ratio, and S be the state space set. Action: For the resource scheduling problem in a hybrid service coexistence system, each agent decides to use arbitrary sub-bandwidth for transmission and the transmission power within that sub-bandwidth; considering the downlink transmission power of F-AP, the transmission power control for mMTC services is defined as P. mMTC ={50,100,150,200}mW; therefore, the dimension of the action space is 4*H, and each action corresponds to a sub-bandwidth and transmission power selection; S32: Solve for the optimal policy in the multi-intelligence reinforcement learning model, and implement the resource scheduling policy with optimal throughput under mixed business by designing a suitable reward function and error function; Reward function: Each agent adjusts its scheduling strategy based on the reward value obtained from each decision, thereby gradually approaching the optimal decision; the goal is to improve network throughput performance and mMTC service transmission success rate in mixed service scenarios within time interval T. The reward function for the mMTC service communication link is set as follows: Among them, C max For Shannon capacity, C re-mMTC λ represents the remaining mMTC traffic volume; μ1 is the URLLC communication link agent factor, indicating whether there is a URLLC service to be transmitted within the current sub-bandwidth; μ2 is the mMTC communication link agent factor, indicating whether there is an mMTC service to be transmitted within the current sub-bandwidth; λ is a hyperparameter used to adjust the network's willingness to access mMTC services. The reward function for the entire model is set as follows: Where, ω URLLC ω is a weighting factor used to adjust the transmission volume of URLLC services. mMTC This is a weighting factor used to adjust the transmission volume of mMTC services; The error function is defined as follows: Where γ is the loss factor, s t Let a be the state at time t. t Let A be the action at time t, and let A be the set of actions, θ be the action at time t. t Let θ be the parameter before updating in the stochastic gradient descent method. t 'Q represents the updated parameters in the stochastic gradient descent method.' t The empirical value at time t.

2. The resource scheduling method according to claim 1, characterized in that, Step S1 specifically includes the following steps: S11: In a fog cell, all IoT tasks generated by IoT devices are collected by the scheduler. The F-AP transmits the channel information collected in the previous cycle to the scheduler. The scheduler combines this channel information to allocate the current IoT task transmission queue and make resource allocation decisions for each IoT task. S12: Defined in a fog cell, the fog access node F-AP transmits H URLLC services, and there are J mMTC services within the coverage area of ​​the F-AP, which are communication between IoT devices; among them, the URLLC service obtains high transmission rate service by connecting with the F-AP, while the mMTC service is the information exchange between IoT devices; assuming that all devices use a single antenna in the scenario, the set of URLLC services in the scenario is represented as H={0,…,h}, and the set of mMTC services is represented as J={0,…,j}; S13: Assuming the channel fading is the same within a sub-bandwidth and different sub-bandwidths are independent of each other, then within a time slot interval; the throughput C of URLLC service in the network. URLLC Represented as: in, B0 is the code length occupied by the URLLC service, T is the transmission duration, C is the channel capacity per unit length in infinite code length (i.e., the Shannon limit), V is the channel dispersion, and Q is the channel length of the URLLC service. -1 (·) is a function The inverse function of , where ε is the bit error rate; The throughput C of mMTC services in the network mMTC Represented as: in, Code length occupied by mMTC services.

3. The resource scheduling method according to claim 2, characterized in that, Step S2 specifically includes the following steps: S21: While satisfying the requirements of low latency and high reliability transmission in URLLC communication links, construct an mMTC service access control model; Among them, C j,t C represents the communication link capacity of the j-th mMTC service within time slot t; mMTC,t The amount of traffic that the mMTC service needs to transmit within time slot t, where n is the sub-bandwidth sequence number, N is the total number of sub-bandwidths, and ρ j,h When the h-th URLLC service and the j-th mMTC service are in the same sub-bandwidth, the mMTC service bandwidth allocation coefficient is used. S22: Combining the data characteristics of different services and jointly considering the transmission power control of different services, a throughput optimization model for mixed services in fog cells is constructed: Where, n max P represents the maximum number of symbols that can be used within a time slot. h,max P represents the maximum transmission power for all URLLC services within the total bandwidth. j,max P represents the maximum transmission power of all mMTC services within the total bandwidth. h,n The power required to transmit the h-th URLLC service within the n-th sub-bandwidth; P j,n The power required to transmit the j-th mMTC service within the n-th sub-bandwidth; ρ j,n Assign control coefficients to sub-bandwidths, ρ j,n ∈{0,1}, when ρ j,n When ρ = 1, it means that the j-th mMTC service can use the n-th sub-bandwidth for transmission. j,n =0 indicates that the nth sub-bandwidth cannot be used for transmission; ρ j',n Assign control coefficients to the sub-bandwidths of the remaining mMTC services, j'∈J, and j'≠j; if ρ j',n =1 indicates that more mMTC services can be added to the current sub-bandwidth for transmission.

4. The resource scheduling method according to claim 3, characterized in that, Step S3 specifically includes: modeling the resource scheduling problem of different services as a multi-agent reinforcement learning problem; and using reinforcement learning to find the resource scheduling strategy with optimal throughput under different environments through an iterative approach; wherein, each communication link between URLLC service and mMTC service is regarded as an agent, which selects an appropriate scheduling strategy to execute in the current environment by collecting experience from different decisions in different environments; when multiple services need to be transmitted simultaneously, multiple communication links form a multi-agent cluster to jointly explore the sub-bandwidth scheduling strategy and power control strategy in the current environment.

Citation Information

Patent Citations

  • Dynamic resource allocation method for high-concurrency multi-service industrial 5G network

    CN111629380A

  • Low-orbit satellite hopping beam optimization method based on migration deep reinforcement learning

    CN114362810A