OFDMA and mu-mimo joint scheduling method for dynamic time-sensitive service deterministic delay guarantee
By constructing a hierarchical deep reinforcement learning architecture and utilizing the collaborative optimization of resource allocation between the main agent and sub-agents, the problem of insufficient resource utilization in mixed service scenarios of wireless communication systems is solved, achieving deterministic latency guarantee for time-sensitive streams and throughput improvement for best-effort streams.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-03-10
- Publication Date
- 2026-06-12
Smart Images

Figure CN122205633A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to resource management and scheduling technology in wireless communication systems, specifically to a joint scheduling method of OFDMA and MU-MIMO for deterministic delay guarantee of dynamic time-sensitive services. Background Technology
[0002] In recent years, network demands for applications such as industrial control, robot collaboration, and vehicular communication have become increasingly diversified. These include both time-sensitive services with stringent requirements for end-to-end latency determinism and reliability, and best-effort high-capacity services with throughput as the primary objective. To support these mixed service scenarios, Wi-Fi technologies, represented by the IEEE 802.11 series, have introduced key technologies such as Orthogonal Frequency Division Multiple Access (OFDMA) and Multiple-User Multiple-Input Multiple-Output (MU-MIMO) in Wi-Fi 6 / 6E (802.11ax) and its subsequent evolution to 802.11be. These technologies enable access points to achieve finer-grained resource allocation and parallel scheduling across time, frequency, and spatial dimensions, thereby significantly improving system concurrent throughput and reducing average latency. However, existing solutions still face significant challenges in deterministic performance metrics such as worst-case latency, jitter, and reliability, necessitating the introduction of more deterministic guarantee strategies in scheduling mechanisms and enhanced cross-layer collaborative optimization.
[0003] Existing research largely focuses on improving system throughput or optimizing average latency, lacking in-depth exploration of deterministic latency guarantee mechanisms required for time-sensitive services. Although some studies have addressed the application of OFDMA in time-sensitive networks, these are typically focused on single service types and fail to systematically resolve the conflicts and coordination issues in resource allocation between time-sensitive flows (TS / TT) and best-effort flows (BE) in mixed service scenarios. In such heterogeneous service networks, existing methods still have significant limitations in fully utilizing the combined resources of OFDMA and MU-MIMO to maximize the throughput of best-effort services while ensuring latency and reliability constraints for time-sensitive flows.
[0004] In recent years, reinforcement learning, especially deep reinforcement learning, has been widely applied in the field of wireless resource scheduling. Among them, the master-agent-sub-agent architecture has shown good potential in optimizing resource allocation. However, existing deep reinforcement learning-based methods mostly focus on improving system throughput or optimizing statistical latency performance, lacking specific guarantee mechanisms for the deterministic requirements of time-sensitive flows. Especially in the combined scheduling mode of OFDMA and MU-MIMO, how to effectively improve overall throughput while meeting the strict latency constraints of time-sensitive services remains a key challenge that urgently needs to be addressed. Summary of the Invention
[0005] To address the above problems, this invention provides a joint scheduling method of OFDMA and MU-MIMO for deterministic delay guarantees for dynamic time-sensitive services, comprising the following steps:
[0006] S1. Construct a system model for a wireless communication downlink network, which includes one access point with L antennas and K sites equipped with single antennas; wherein, the time-frequency resources of this system model are divided into several resource blocks RU(m,n,t), where m represents the frequency domain specification level, n represents the resource block index within the specification level m, and t represents the scheduling time slot index; the maximum number of spatial streams supported by each resource block RU(m,n,t) is... Determined by the following formula:
[0007] ,
[0008] In the formula, This indicates the number of antennas at the access point. Indicates the number of antennas at the site. This indicates a round-down operation;
[0009] The system model includes PTS users and BE users, defining each PTS user. Flow characteristics f i for , Indicates the flow period. Indicates the data size. Indicates the timestamp of the data packet arriving at the access point. Indicates the deadline for data packets. Indicates the allowable jitter range. Let represent the set of PTS users; each PTS user satisfies deterministic delay constraints.
[0010] S2. Construct a hierarchical deep reinforcement learning architecture consisting of one master agent and several sub-agents; wherein, the master agent determines a resource block combination based on the global state and performs pre-scheduling on the resource block combination through the PTS user priority reservation mechanism; subsequently, each sub-agent performs subsequent scheduling on BE users in the remaining space of its corresponding resource block; the pre-scheduling result is merged with the subsequent scheduling results of all sub-agents to obtain the allocation mapping;
[0011] S3. Construct a comprehensive reward function and jointly train the main agent and sub-agents;
[0012] S4. Utilize the trained hierarchical deep reinforcement learning architecture to generate a resource allocation scheme for each scheduling slot.
[0013] The beneficial effects of this invention are:
[0014] The hierarchical hybrid scheduling framework proposed in this invention clearly divides the work into two levels: macro-level resource combination selection and micro-level serialized user selection within resource blocks. It achieves collaborative optimization through joint training, fundamentally reducing the combinatorial search space and improving the feasibility of real-time decision-making. By using statistical summary-type global state at the main agent layer and fine-grained per-user channel characteristics at the sub-agent layer, it achieves a balance between decision-making generalizability and local refinement capabilities. The introduction of PTS reliability determination and priority reservation mechanisms based on redundant sample counts ensures end-to-end deterministic reliability even in scenarios with limited physical layer resources and complex spatial correlations among users. While guaranteeing PTS service, it improves the spectrum utilization and throughput performance of BE services by using a hybrid approach of OFDMA and MU-MIMO and dynamically switching multiplexing modes at the main agent layer. Simultaneously, it suppresses resource abuse through fairness and redundancy penalty terms, balancing system throughput and user fairness. Attached Figure Description
[0015] Figure 1 This is an overall flowchart of the method provided in the embodiments of the present invention;
[0016] Figure 2 This is a schematic diagram of the main agent action space provided in an embodiment of the present invention;
[0017] Figure 3 This is a diagram illustrating the hierarchical deep reinforcement learning agent framework provided in an embodiment of the present invention.
[0018] Figure 4 The hierarchical deep reinforcement learning flowchart provided for embodiments of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Please see Figures 1-4 This invention provides a joint scheduling method of OFDMA and MU-MIMO for deterministic delay guarantee of dynamic time-sensitive services, including the following steps:
[0021] S1. Construct a system model for the wireless communication downlink network.
[0022] This invention provides a system model based on the IEEE 802.11ax downlink scenario, comprising one access point (AP) with L antennas and K stations (STAs) equipped with a single antenna. The time-frequency resources of this system model are divided into several resource blocks RU(m,n,t), where m represents the frequency domain specification level. This represents the resource block index within the frequency domain specification level m, and t represents the scheduling slot index; specifically... Classifying resource block bandwidth specifications at the frequency domain level can characterize different resource block bandwidth specifications such as 242-tone and 106-tone. Used to distinguish different resource blocks under the same frequency domain specification level.
[0023] The maximum number of spatial flows supported by each resource block RU(m,n,t) Determined by the following formula:
[0024] ,
[0025] In the formula, This indicates the number of antennas at the access point. Indicates the number of antennas at the site. This indicates a round-down operation.
[0026] In this embodiment of the invention, the channel quality between the AP and STA is measured and evaluated based on the actual environment model.
[0027] In this embodiment of the invention, to distinguish between time-sensitive flows and best-effort flows in data transmission, the system model defines two user types: PTS users with latency guarantees and BE users employing a best-effort service strategy. To avoid overly complex system models, it is assumed that each station generates only one type of service within the same scheduling time slot.
[0028] In some embodiments, the system model employs a combination of actual measurements and theoretical analysis to construct a downlink multi-user channel model, thereby characterizing the channel quality between the access point and the site.
[0029] Specifically, the downlink multi-user channel model defines the first... The received signal for each user consists of three parts: the desired signal, co-channel interference from other sites, and additive white Gaussian noise; among which, Indicates the first The channel vector of each user on resource block RU(m,n,t) is used to reflect channel state information; Indicates the first A user's precoded vector on resource block RU(m,n,t) is used for beamforming; Indicates the first The signal-to-interference-plus-noise ratio (SIR) of a user on resource block RU(m,n,t) is determined by the expected signal power, co-channel interference power, and noise power. It can represent the total set of users or the set of sites.
[0030] Define indicator variables When the first When a user is scheduled to resource block RU(m,n,t), =1, otherwise =0. Furthermore, the first... The transmit power and data symbols for each user are respectively represented as follows: and .
[0031] In a practical system, the achievable data rate for each user is mapped using a predefined modulation and coding scheme (MCS) lookup table, which establishes the correspondence between signal-to-interference-plus-noise ratio (SINR) and spectral efficiency. The system selects the highest-order MCS scheme that meets the SINR requirements and is feasible for each user, thereby maximizing the transmission rate.
[0032] In some embodiments, to address the specific needs of time-sensitive services, the first... Traffic characteristics of a PTS user f i for:
[0033] ,
[0034] Indicates the flow period. Indicates the data size. Indicates the timestamp of the data packet arriving at the current node. Indicates the deadline for data packets. Indicates the allowable jitter range. This represents the set of PTS users; each PTS user satisfies a deterministic delay constraint. .
[0035] Specifically, the basic time slot of the system model is taken as the greatest common divisor of the traffic cycles of all PTS users, denoted as . A supercycle is the least common multiple of the traffic cycles of all PTS users, denoted as . .
[0036] By systematically analyzing the various delay components in the end-to-end transmission process, deterministic delay constraints can be defined as follows:
[0037]
[0038] In the formula, Indicates the first The processing latency for each PTS user includes the time consumed in the steps required for the access point to process data frames, such as decapsulation / encapsulation, MAC / PHY protocol stack operations, encryption / decryption, encoding / decoding, etc. Indicates the first The queuing delay for each PTS user is the time interval from the arrival of a data packet to its successful scheduling. Indicates the first The propagation delay for each PTS user is proportional to the physical distance between the AP and STA. In indoor environments, this value is typically calculated based on a light-speed propagation model. Indicates the first The data rate that a PTS user can reach in the current resource block.
[0039] In some embodiments, it is assumed that the first The data packets arriving at each BE user follow an average rate of... The Poisson process, whose queue backlog in any scheduling time slot t is denoted as . The performance metrics of the BE service are determined by the observation window. Cumulative effective throughput obtained within Decide:
[0040]
[0041] In the formula, Indicates the first The transmission rate of each BE user on the current resource block The basic time slots representing the system model, This represents the set of BE users.
[0042] S2. Construct a hierarchical deep reinforcement learning architecture consisting of one master agent and several sub-agents; wherein, the master agent determines a resource block combination based on the global state and performs pre-scheduling on the resource block combination through the PTS user priority reservation mechanism; subsequently, each sub-agent performs subsequent scheduling on BE users in the remaining space of its corresponding resource block; the pre-scheduling result is merged with the subsequent scheduling results of all sub-agents to obtain the allocation mapping.
[0043] In some embodiments, the master agent receives the global state from the radio environment in each scheduling time slot t. Then, the global state is processed through a multilayer perceptron (MLP). Output a resource block combination ,like Figure 3 As shown, the primary agent's decision determines which resource blocks and their spatial flow configurations are activated.
[0044] The main agent's specific processing procedures include:
[0045] S211. Constructing the global state , This represents the normalized total amount of data to be transmitted across the entire network. Indicates global static evaluation features, This represents the static evaluation characteristics of the q-th resource block (q=1,2,…,Q), where Q represents the number of resource blocks. This represents a summary statistical indicator of inter-user channel spatial correlation calculated based on the current activity layout. Static evaluation characteristics of each resource block include whether it supports spatial reuse and the maximum number of spatial streams it can carry. The theoretical throughput estimate is calculated based on bandwidth and modulation / coding scheme. Inter-user channel spatial correlation summary statistics include the mean, variance, skewness, and kurtosis of correlation statistics, as well as several graph structure summary indicators, such as the number of connected components, average degree, and clustering coefficient.
[0046] S212. The main agent, based on the global state... Select a resource block combination from the pre-defined resource block combination table. ,include:
[0047] S2121. Divide all resource block combinations in the predefined resource block combination table into a spatially reusable subset and a non-spatially reusable subset, which are disjoint.
[0048] In some embodiments, the definition of a predefined resource block combination is as follows: Figure 2 As shown in the figure, this example illustrates 30 different combinations of resource blocks, each specifying the number of resource blocks it contains and the corresponding spatial flow configuration for each resource block.
[0049] S2122. Change the global state Input to the encoding network to obtain hidden vectors The encoding network employs a fully connected layer with an output dimension of 256, and a ReLU layer is cascaded after the fully connected layer.
[0050] S223. Hidden vector Output resource block combinations through a gating decision network. ;in:
[0051] The gating decision network consists of a fully connected layer with an output dimension of 256, followed by parallel connections to a reusable combinatorial network, a non-reusable combinatorial network, and a gating network. Both the reusable and non-reusable combinatorial networks are composed of a fully connected layer, with output dimensions of the number of spatially reusable resource block combinations and the number of non-spatially reusable resource block combinations, respectively. The gating network employs a fully connected layer with an output dimension of 2 and a softmax layer with a temperature parameter. The gating network determines whether to enable spatial reuse mode, the reusable combinatorial network calculates the selection probability of each resource block combination in the spatially reusable subset, and the non-reusable combinatorial network calculates the selection probability of each resource block combination in the non-spatially reusable subset.
[0052] Based on the outputs of reusable combinatorial networks, non-reusable combinatorial networks, and gated networks, a resource block combination is determined from a predetermined resource block combination table through combinatorial analysis, denoted as […]. Resource block combination The specific resource blocks involved and their corresponding spatial flow configurations will be activated.
[0053] Specifically, if the gating result is to enable spatial reuse mode, the resource block combination corresponding to the maximum selection probability of the reusable combinatorial network output is selected for activation; if the gating result is to disable spatial reuse mode, the resource block combination corresponding to the maximum selection probability of the non-reusable combinatorial network output is selected for activation.
[0054] Through the above mechanism, the main agent can prioritize space reuse in low-interference scenarios to improve system concurrency throughput; in high-correlation (strong interference) scenarios, it can automatically switch to non-reuse mode to ensure the performance and reliability of deterministic services.
[0055] In some embodiments, each sub-agent is associated with its resource block. The process of performing BE user selection includes:
[0056] S221. Obtain the local observation vector of the resource block r associated with the sub-agent. Simultaneously, obtain the channel feature matrix of all BE users in the resource blocks associated with the sub-agent. The local observation vector This includes normalized bandwidth, a flag indicating whether spatial multiplexing is supported, the maximum number of available spatial streams (excluding the number of spatial streams reserved for PTS users in pre-scheduling), and a theoretical throughput estimate. Channel feature matrix. The feature vector of each BE user is obtained by processing the channel state information (if a complex baseband matrix exists, the real and imaginary parts are retained and concatenated in parallel, or the dimensionality is reduced by feature extraction methods).
[0057] S222. Transfer the local observation vector Mapped to RU embedding vectors via resource block encoder At the same time, the channel feature matrix Each row element is mapped through a user encoder to obtain a corresponding feature embedding vector, resulting in a user embedding matrix. .
[0058] Specifically, the resource block encoder uses two fully connected layers with an output dimension of 128, and each fully connected layer is cascaded with a ReLU layer. The user encoder uses two fully connected layers with an output dimension of 256, and each fully connected layer is cascaded with a ReLU layer.
[0059] S223. Based on RU embedding vectors With user embedding matrix Through a serialized decision-making mechanism, targets are iteratively selected from all BE users in the system model, ultimately generating a subset of BE users; in the k-th iteration, the selection process includes:
[0060] Obtain the feature embedding vector of the BE user selected in the (k-1)th iteration. and associate it with the RU embedding vector The concatenated data is then input into a recurrent neural network (RNN) to obtain the hidden states. :
[0061]
[0062] In the formula, Let represent the hidden state during the selection process in the (k-1)th iteration. Specifically, the initial hidden state. RU embedding vector In the first iteration, since no BE users have been selected yet, at this time... It is set as an all-zero vector to initiate the iterative selection process.
[0063] Based on the hidden state With user embedding matrix Calculate the original priority score for each BE user;
[0064] Based on the original priority scores, the selection probability of each BE user is calculated and expressed as follows:
[0065]
[0066] In the formula, Indicates the first The probability of a BE user's choice. In the k-th iteration, the... The original priority score of each BE user Represents the set of BE users; Indicates the first The availability mask for each BE user; if the BE user has already been selected, is in an empty queue, or does not meet the quota priority, then... =0, otherwise =1.
[0067] The user with the highest selection probability (BE) is selected and added to the user subset.
[0068] Repeat the selection process above until the number of selected users reaches the maximum available space stream count for the resource block or there are no more qualified candidate users.
[0069] It should be noted that if the resource block associated with each sub-agent has no available space flow, the above scheduling process will not be performed.
[0070] In some embodiments, pre-scheduling of resource block combinations is performed through the PTS user-priority reservation mechanism, including:
[0071] S231. Calculate the urgency level for each PTS user with data to be transmitted. , represented as:
[0072]
[0073] In the formula, Indicates the current time.
[0074] S232. Based on the order of urgency from high to low, reserve space flow for PTS users in turn on the currently activated resource blocks with available space flow, until their minimum necessary redundancy requirements are met or available resources are exhausted; in addition, a correlation check is performed during the reservation process: calculate the similarity between the PTS user to be reserved and the set of users already reserved in the current resource block, and if the similarity exceeds a preset correlation threshold, skip the current resource block.
[0075] After the PTS user's priority reservation is completed, the sub-agent only performs BE user scheduling on the remaining space stream resources, thereby ensuring the deterministic latency requirements of time-sensitive streams and improving the overall throughput performance of non-time-sensitive streams.
[0076] S3. Construct a comprehensive reward function and jointly train the main agent and sub-agents.
[0077] In some embodiments, the combined reward function is used to simultaneously optimize the deterministic latency and reliability of time-sensitive flows, as well as the throughput of best-effort flows, and is expressed as:
[0078]
[0079] In the formula, This represents the BE service throughput benefit term, used to measure the average throughput satisfaction of non-time-sensitive flows within the current scheduling cycle; Represents the reliability gain of the i-th PTS user, used to incentivize time-sensitive stream redundancy transmission behavior that meets its application-layer reliability requirements; This represents the redundancy resource constraint penalty for the i-th PTS user, used to penalize excessive resource allocation that exceeds the minimum redundancy required to meet the reliability threshold; This represents a fairness adjustment term, used to penalize the uneven reliability gain among PTS users and the uneven throughput distribution among BE users, thereby prompting the post-trained resource allocation strategy to maintain the reliability of deterministic services while taking into account the service fairness of all users. , , , This represents the weighting coefficient.
[0080] Specifically, the BE service throughput benefit is calculated using the actual achievable throughput of each BE user on each allocated resource block, and then normalized to characterize the improvement in throughput performance relative to the expected level. The reliability gain for each PTS user is the probability of successful redundant transmission. PTS business target reliability threshold The difference, p is the single transmission success rate, and n1 is the number of redundant samples allocated to PTS users within this scheduling period; The fairness adjustment term is quantified based on the extent to which the actual resource allocation to PTS users exceeds their minimum redundancy. The value increases as the redundancy allocation exceeds the minimum redundancy, to avoid consuming excessive space flow resources and affecting overall scheduling efficiency. The fairness adjustment term is constructed based on the Jain fairness index.
[0081] Specifically, the minimum redundancy is defined as the minimum number of redundant samples that satisfies the following conditions. :
[0082]
[0083] Minimum number of redundant samples This represents the number of independently decodeable redundant transmission replicas allocated to a PTS user within the same scheduling period under a hybrid OFDMA and MU-MIMO resource block structure, used to ensure that its end-to-end success rate is not lower than the aforementioned reliability target.
[0084] In some embodiments, a framework of centralized training and distributed execution is used to jointly train the main agent and sub-agents. For example... Figure 4 As shown, during the training phase, the main agent Actor, sub-agent Actor, and global Critic form a closed-loop training architecture. The main agent and sub-agents perform corresponding actions in the wireless environment, and the environment calculates scores based on the comprehensive reward function and feeds back state transitions (…). ) and instant comprehensive rewards All interaction data is stored uniformly in the experience buffer and organized into trajectory sample vectors:
[0085]
[0086] in, This represents the merged state vector, which contains both the global summary information required by the main agent and the local features available to the sub-agents. This indicates a joint action, which is composed of the combination of resource blocks selected by the main agent and the user selection sequence within each resource block by the sub-agents. Represents the logarithmic probability of an action. This is the value network's estimate of the current state. The training and update module reads samples in batches from the experience replay buffer, performs value regression on the global Critic network, and synchronously updates the policy parameters of the main agent and sub-agents based on Proximal Policy Optimization (PPO). During the execution phase, the main agent and sub-agents no longer access the global state, but only complete inference and scheduling based on their own local observations, realizing a structure of "centralized training and distributed execution".
[0087] As the core policy update algorithm, PPO's pruning objective can be simply written as:
[0088]
[0089] in The probability ratio represents the ratio of the probability of the new and old strategies choosing the same action. The advantage function is used to evaluate the performance in state 1. Select action The degree of advantage relative to the average level; ϵ is the pruning hyperparameter, with Figure 4For reference, the key points of the training closed loop are: the environment generates samples → samples are put into the buffer → the buffer provides training batches to the Critic and optimizer → the optimizer calculates the advantage and updates the parameters based on the batches → the parameters are sent back to the Actor, and the Critic's estimated value is written back to the buffer for subsequent regression and advantage calculation.
[0090] S4. Utilize the trained hierarchical deep reinforcement learning architecture to generate resource allocation schemes for each scheduling slot.
[0091] After the model is deployed, the current network status is observed in each scheduling time slot, and a resource allocation mapping map is generated through forward computation. The allocation results are sent to the physical layer for execution through a standard interface to complete the hybrid resource allocation of OFDMA and MU-MIMO.
[0092] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A joint scheduling method of OFDMA and MU-MIMO for deterministic delay guarantee of dynamic time-sensitive services, characterized in that, Includes the following steps: S1. Construct a system model for a wireless communication downlink network, which includes one access point configured with L antennas and K sites equipped with single antennas; wherein, the time-frequency resources of this system model are divided into several resource blocks RU(m,n,t), where m represents the frequency domain specification level. The index of the resource block within the frequency domain specification level m represents the index of the scheduling slot; t represents the index of the scheduling slot; the maximum number of spatial streams supported by each resource block RU(m,n,t). Determined by the following formula: , In the formula, This indicates the number of antennas at the access point. Indicates the number of antennas at the site. This indicates a round-down operation; The system model includes PTS users and BE users, defining the first... Traffic characteristics of a PTS user f i for , Indicates the flow period. Indicates the data size. Indicates the timestamp of the data packet arriving at the access point. Indicates the deadline for data packets. Indicates the allowable jitter range. Let represent the set of PTS users; each PTS user satisfies deterministic delay constraints. S2. Construct a hierarchical deep reinforcement learning architecture consisting of one master agent and several sub-agents; wherein, the master agent determines a resource block combination based on the global state and performs pre-scheduling on the resource block combination through the PTS user priority reservation mechanism; subsequently, each sub-agent performs subsequent scheduling on BE users in the remaining space of its corresponding resource block; the pre-scheduling result is merged with the subsequent scheduling results of all sub-agents to obtain the allocation mapping; S3. Construct a comprehensive reward function and jointly train the main agent and sub-agents; S4. Utilize the trained hierarchical deep reinforcement learning architecture to generate a resource allocation scheme for each scheduling slot.
2. The OFDMA and MU-MIMO joint scheduling method for deterministic delay guarantee of dynamic time-sensitive services according to claim 1, characterized in that, The system model also includes a downlink multi-user channel model, in which: Definition of the first The received signal for each user consists of three parts: the desired signal, co-channel interference from other sites, and additive white Gaussian noise; among which, Indicates the first The channel vectors of each user on resource block RU(m,n,t) Indicates the first The precoded vectors of each user on resource block RU(m,n,t) Indicates the first The signal-to-interference-plus-noise ratio (SIR) of each user on resource block RU(m,n,t); Represents the total set of users; Define indicator variables When the first When a user is scheduled to resource block RU(m,n,t), =1, otherwise =0.
3. The OFDMA and MU-MIMO joint scheduling method for deterministic delay guarantee of dynamic time-sensitive services according to claim 1, characterized in that, The deterministic delay constraint is expressed as: , In the formula, Indicates the first Processing latency for each PTS user Indicates the first Queuing delay for PTS users Indicates the first Propagation delay for each PTS user Indicates the first The data rate that a PTS user can reach in the current resource block.
4. The OFDMA and MU-MIMO joint scheduling method for deterministic delay guarantee of dynamic time-sensitive services according to claim 1, characterized in that, The main agent adjusts the global state in each scheduling time slot t. Determine the combination of resource blocks include: S211. Constructing the global state , This represents the normalized total amount of data to be transmitted across the entire network. Indicates global static evaluation features, This represents the static evaluation characteristics of the q-th resource block (q=1,2,…,Q), where Q represents the number of resource blocks. This represents a summary statistical indicator of channel spatial correlation between users; the static evaluation characteristics of each resource block include whether spatial reuse is supported, the maximum number of spatial streams, and the theoretical throughput estimate. S212. The main agent, based on the global state... Select a resource block combination from the pre-defined resource block combination table, including: S2121. Divide all resource block combinations in the predefined resource block combination table into a spatially reusable subset and a non-spatially reusable subset, which are disjoint. S2122. Change the global state Input to the encoding network to obtain hidden vectors The encoding network employs a fully connected layer with an output dimension of 256, and a ReLU layer is cascaded after the fully connected layer. S2123. Hidden vector Output resource block combinations through a gating decision network. ;in: The gating decision network consists of a fully connected layer with an output dimension of 256, followed by parallel connections to a reusable combinatorial network, a non-reusable combinatorial network, and a gating network. The gating network decides whether to enable the spatial reuse mode, the reusable combinatorial network calculates the selection probability of each resource block combination in the spatially reusable subset, and the non-reusable combinatorial network calculates the selection probability of each resource block combination in the non-spatially reusable subset. Based on the outputs of reusable combinatorial networks, non-reusable combinatorial networks, and gated networks, a resource block combination is determined from a predetermined resource block combination table through combinatorial analysis, denoted as . Resource block combination The specific resource blocks involved and their corresponding spatial flow configurations will be activated.
5. The OFDMA and MU-MIMO joint scheduling method for deterministic delay guarantee of dynamic time-sensitive services according to claim 1, characterized in that, The scheduling process for each sub-agent includes: S221. Obtain the local observation vector of the resource block associated with the sub-agent. Simultaneously, obtain the channel feature matrix of all BE users in the resource blocks associated with the sub-agent. The local observation vector This includes normalized bandwidth, whether spatial multiplexing is supported, maximum available spatial stream count, and theoretical throughput estimate; S222. Transfer the local observation vector Mapped to RU embedding vectors via resource block encoder At the same time, the channel feature matrix Each row of elements is mapped through a user encoder to obtain a user embedding matrix. ; S223. Based on RU embedding vectors With user embedding matrix Through a serialized decision-making mechanism, targets are iteratively selected from all BE users in the system model, ultimately generating a subset of BE users; in the k-th iteration, the selection process includes: Obtain the feature embedding vector of the BE user selected in the (k-1)th iteration. and associate it with the RU embedding vector The concatenated data is then input into a recurrent neural network to obtain the hidden state. : , In the formula, This represents the hidden state during the selection process in the (k-1)th iteration; Based on the hidden state With user embedding matrix Calculate the original priority score for each BE user; Based on the original priority score and the availability mask, the probability of selection for each BE user is calculated and expressed as follows: , In the formula, Indicates the first The probability of a BE user's choice. In the k-th iteration, the... The original priority score of each BE user Represents the set of BE users; Indicates the first The availability mask for each BE user; if the BE user has already been selected, is in an empty queue, or does not meet the quota priority, then... =0, otherwise =1; The user with the highest selection probability (BE) is selected and added to the user subset.
6. The OFDMA and MU-MIMO joint scheduling method for deterministic delay guarantee of dynamic time-sensitive services according to claim 1, characterized in that, The resource block combination is pre-scheduled through the PTS user priority reservation mechanism, including: S231. Calculate the urgency level for each PTS user with data to be transmitted. ,in, Indicates the current time; S232. Based on the order of urgency from high to low, reserve space flow for PTS users in turn on the currently activated resource blocks with available space flow, until their minimum necessary redundancy requirements are met or available resources are exhausted; perform a correlation check during the reservation process: calculate the similarity between the PTS user to be reserved and the set of users already reserved in the current resource block, and if the similarity exceeds the preset correlation threshold, skip the current resource block.
7. The OFDMA and MU-MIMO joint scheduling method for deterministic delay guarantee of dynamic time-sensitive services according to claim 1, characterized in that, The comprehensive reward function Re is expressed as: , In the formula, This represents the BE (Business Entity) throughput benefit item. This represents the reliability gain for the i-th PTS user. This represents the redundancy resource constraint penalty for the i-th PTS user. Indicates the fairness adjustment item, , , , This represents the weighting coefficient.