Machine learning-based resource block scheduling in multi-user multiple-input multiple-output network

The proposed RB scheduling apparatus uses a neural network to combine precoders and apply SVD for UE scheduling decisions, addressing the limitations of existing ML-based schedulers by optimizing RB or RBG allocation based on spatial diversity and ranks, enhancing scheduling efficiency and accuracy.

WO2025171858A1PCT designated stage Publication Date: 2025-08-21NOKIA SOLUTIONS & NETWORKS OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/053512
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-12
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing ML-based RB schedulers for MU-MIMO networks are limited in their ability to account for rank adaptation, leading to suboptimal performance in scheduling UEs, particularly for rank-1 MU-MIMO scenarios.

Method used

An RB scheduling apparatus and method that utilizes a neural network (NN) to combine precoders of scheduled and candidate UEs, apply Singular Value Decomposition (SVD) to determine singular values, and make scheduling decisions based on spatial diversity and ranks, considering frequency-specific information to optimize RB or RBG allocation.

Benefits of technology

Enables efficient RB or RBG scheduling in MU-MIMO networks with low computational complexity, accounting for UE spatial diversity and ranks, improving scheduling accuracy and reducing execution time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024053512_21082025_PF_FP_ABST
    Figure EP2024053512_21082025_PF_FP_ABST
Patent Text Reader

Abstract

A technical solution is provided, which enables Resource Block (RB) scheduling for UEs in a Multi-User Multiple-Input Multiple-Output (MU-MIMO) network based on the spatial diversity and ranks of the UEs. For this purpose, one or more scheduled UEs already having one or more RBs for a Transmission Time Interval (TTI) and one or more candidate UEs having no RB for the TTI are determined. For each candidate UE, its precoder is combined with the precoder of each scheduled UE. The combined set of precoders is subjected to Singular Value Decomposition (SVD), whereupon one or more singular values are selected based on the rank sum of the candidate UE and each scheduled UE. The decision whether to schedule the RB(s) for the candidate UE is made using at least one NN receiving the rank sum and the singular value(s) (or a value derived from the singular value(s)) as input.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] MACHINE LEARNING-BASED RESOURCE BLOCK SCHEDULING IN MULTI-USER MULTIPLE-

[0002] INPUT MULTIPLE-OUTPUT NETWORK

[0003] TECHNICAL FIELD

[0004] The present disclosure relates generally to the field of wireless communications. In particular, the present disclosure relates to a Machine Learning (ML)-based Resource Block (RB) scheduling apparatus and method in a Multi-User Multiple-Input Multiple-Output (MU- MIMO) network, as well as to a corresponding computer program product.

[0005] BACKGROUND

[0006] It has been recently shown that a properly configured ML model can learn to efficiently allocate RBs (i.e., time-frequency resources) or RB Groups (RBGs) (each typically including 4- 16 continuous RBs) to User Equipments (UEs). In this sense, a ML-based RB scheduler can replace the existing RB schedulers relying on heuristic algorithms. More specifically, it has been also shown that the ML-based RB scheduler can outperform the heuristic RB schedulers in spectral efficiency. Furthermore, the ML-based RB scheduler has been proved to dominate in scheduling execution time, by doing real-time RB scheduling decisions for uplink (UL) approximately 70-80% faster than the existing RB schedulers using complex state-of-the-art heuristic uplink scheduling algorithms.

[0007] Despite the above-indicated domination of the ML-based RB scheduler in the spectral efficiency and the real-time execution time, there is room for further improvements of the ML-based RB scheduler. More specifically, although the existing ML-based RB scheduler design is configured to perform RB scheduling in a frequency domain and learn UE pairing for frequency selective MU-MIMO scheduling, it can only do all this very well for rank-1 MU- MIMO. In other words, the existing ML-based RB scheduler design is not configured to take rank adaptation into account when performing RB scheduling for UEs of interest. SUMMARY

[0008] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features of the present disclosure, nor is it intended to be used to limit the scope of the present disclosure.

[0009] It is an objective of the present disclosure to provide a technical solution that enables RB or RBG scheduling for UEs in a MU-MIMO network based on the spatial correlation (or diversity) and ranks of the UEs.

[0010] The objective above is achieved by the features of the independent claims in the appended claims. Further embodiments and examples are apparent from the dependent claims, the detailed description, and the accompanying drawings.

[0011] According to a first aspect, an RB scheduling apparatus in a MU-MIMO network is provided. The apparatus comprises at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to operate at least as follows. At first, the apparatus receives initial information comprising: (i) a first list of UEs comprising at least one scheduled UE for which one or more RBs or RBGs have been already scheduled for a target Transmission Time Interval (TTI), (ii) a second list of UEs comprising at least one candidate (i.e., non-scheduled) UE for which said one or more RBs or RBGs are to be scheduled for the target TTI, (iii) a precoder for each of the at least one scheduled UE and the at least one candidate UE, and (iv) a rank for each of the at least one scheduled UE and the at least one candidate UE. The rank indicates a UE-associated number of MU-MIMO layers to be used in the MU-MIMO network. After the initial information is received, the apparatus performs the following operations for each of the at least one candidate UE: obtaining a combined set of precoders by combining the precoder for the candidate UE with the precoder for each of the at least one scheduled UE; obtaining an array of singular values by applying Singular Value Decomposition (SVD) to the combined set of precoders; calculating a sum of ranks by summing the rank for the candidate UE with the rank for each of the at least one scheduled UE; based on the sum of ranks, selecting one or more target singular values from the array of singular values; and obtaining a scheduling decision for the candidate UE with respect to each of said one or more RBs or RBGs for the target TTI by using an ML model. The ML model comprises at least one Neural Network (NN) configured to receive input data comprising: (i) the sum of ranks, and (ii) said one or more target singular values, or a value derived from said one or more target singular values, and output the scheduling decision with respect to each of said one or more RBs or RBGs.

[0012] The apparatus thus configured may take the MU-MIMO aspects (i.e., the spatial diversity of UEs) as well as the different ranks of the UEs in scheduling (both downlink (DL) and UL) RBs or RBGs for the UEs. Furthermore, the apparatus thus configured may perform the frequency and spatial domain RB scheduling with low computational complexity regardless of: a maximum number of UEs to be scheduled to RBs or RBGs, a physical layer numerology / bandwidth (i.e., a number of RBs or RBGs per TTI), (digital / analog / hybrid) conventional and / or ML-based beamformers used in the MU-MIMO network, and a number of Radio Frequency (RF) chains determining a maximum number of spatially scheduled UEs.

[0013] In one example embodiment of the first aspect, the apparatus is further caused to mark the candidate UE as a scheduled UE for at least one RB or RBG of said one or more RBs or RBGs if the scheduling decision with respect to the at least one RB or RBG indicates that the at least one RB or RBG is to be scheduled to the candidate UE. In this way, it is possible to easily and quickly identify those UEs that have not yet been scheduled for a certain RB or RBG.

[0014] In one example embodiment of the first aspect, the at least one NN is trained using one of a Deep Q-Network (DQN) algorithm, a Double DQN (DDQN) algorithm, a Proximal Policy Optimization (PPO) algorithm, and a Soft Actor Critic (SAC) algorithm. By using these reinforcement training algorithms, it is possible to train the at least one NN more efficiently, which may subsequently lead to more accurate prediction of scheduling decisions for UEs.

[0015] In one example embodiment of the first aspect, the at least one NN comprises an input layer of neurons comprising a number of neurons defined by a user-defined maximum number of candidate UEs. In this embodiment, the at least one NN is configured, if a current number of candidates UEs on the second list of UEs is less than the user-defined maximum number of candidate UEs, to pad the input data with one or more zeros such that the input data has a dimension equal to the number of neurons. Since the NN topology is fixed based on the user- defined maximum number of candidate UEs, all leftover NN inputs may be padded with one or more zeros as the number of candidate UEs on the second list of UEs becomes less the user-defined maximum number of candidate UEs during said processing. Thus, also the NN outputs corresponding to the zeroed inputs may be left out from the scheduling-decision making process. For example, if the NN provides probabilities (e.g., Q-values) for each candidate UE, values of invalidated (i.e., for which RB or RBG scheduling has already been performed or is not required at all) UEs on the second list of UEs may be set to minimum (or zero) to avoid their selection. This may allow considering only those UEs which are be needed to be scheduled for the PRBs or RBGs, thereby reducing computational costs even more.

[0016] In one example embodiment of the first aspect, the at least one NN comprises a first NN and a second NN. The first NN is configured to obtain the scheduling decision for each of the at least one candidate UE on the second list of UEs with respect to each of said one or more RBs or RBGs. The second NN is configured to obtain a UE pairing decision for each of the at least one candidate UE on the second list of UEs with respect to each of said one or more RBs or RBGs. The UE pairing decision indicates whetherthe candidate UE needs to share said one or more RBs or RBGs with the at least one scheduled UE on the first list of UEs. With this approach, it is possible to train the first NN to optimize frequency-selective scheduling and, for example, proportional fairness, and the second NN to benefit from the MU -Ml MO aspects without compromising the performance of the UEs scheduled to the same RB or RBG.

[0017] In one example embodiment of the first aspect, the input data for each of the at least one candidate UE on the second list of UEs further comprises frequency-specific information. The frequency-specific information comprises at least one of: wideband Channel State Information (CSI) for the candidate UE, sub-band CSI for the candidate UE, a number of already scheduled RBs or RBGs for the candidate UE, a past average DL throughput for the candidate UE, and buffer status information for the candidate UE. By using this additional frequency-specific information, the NN(s) may predict the scheduling decisions for the candidate UE(s) more accurately.

[0018] In one example embodiment of the first aspect, the at least one NN is trained based on a sequence of tuples by using a reinforcement learning algorithm. Each tuple of the sequence of tuples comprises: the input data for at least one previous candidate UE on the second list of UEs, the scheduling decision outputted by the at least one NN for each of the at least one previous candidate UE, a performance reward resulted from the scheduling decision for each of the at least one previous candidate UE, and the input data for at least one next candidate UE on the second list of UEs. By doing so, it is possible to train the NN(s) for frequency- and spatial-domain RB scheduling in the most efficient way (i.e., with good performance vs complexity characteristics).

[0019] According to a second aspect, an RB scheduling method in a MU -Ml MO network is provided. The method stars with the step of receiving initial information which comprises: (i) a first list of UEs comprising at least one scheduled UE for which one or more RBs or RBGs have been already scheduled for a target TTI, (ii) a second list of UEs comprising at least one candidate UE for which said one or more RBs or RBGs are to be scheduled for the target TTI, (iii) a precoder for each of the at least one scheduled UE and the at least one candidate UE, and (iv) a rank for each of the at least one scheduled UE and the at least one candidate UE. The rank indicates a UE-associated number of MU-MIMO layers to be used in the MU-MIMO network. After that, the method proceeds to the following steps performed for each of the at least one candidate UE: obtaining a combined set of precoders by combining the precoder for the candidate UE with the precoder for each of the at least one scheduled UE; obtaining an array of singular values by applying SVD to the combined set of precoders; calculating a sum of ranks by summing the rank for the candidate UE with the rank for each of the at least one scheduled UE; based on the sum of ranks, selecting one or more target singular values from the array of singular values; and obtaining a scheduling decision for the candidate UE with respect to each of said one or more RBs or RBGs for the target TTI by using a ML model. The ML model comprises at least one NN configured to receive input data comprising: (i) the sum of ranks, and (ii) said one or more target singular values, or a value derived from said one or more target singular values, and output the scheduling decision with respect to each of said one or more RBs or RBGs.

[0020] By using the method according to the second aspect, it is possible to perform RB or RBG scheduling for UEs considering the MU-MIMO aspects (i.e., the spatial diversity of the UEs) as well as the different ranks of the UEs. Furthermore, such frequency and spatial domain RB scheduling may be performed with low computational complexity regardless of: a maximum number of UEs to be scheduled to RBs or RBGs, a physical layer numerology / bandwidth (i.e., a number of RBs or RBGs per TTI), (digital / analog / hybrid) conventional and / or ML-based beamformers used in the MU -Ml MO network, and a number of RF chains determining a maximum number of spatially scheduled UEs.

[0021] In one example embodiment of the second aspect, the method further comprises the step of marking the candidate UE as a scheduled UE for at least one RB or RBG of said one or more RBs or RBGs if the scheduling decision with respect to the at least one RB or RBG indicates that the at least one RB or RBG is to be scheduled to the candidate UE. In this way, it is possible to easily and quickly identify those UEs that have not yet been scheduled for a certain RB or RBG.

[0022] In one example embodiment of the second aspect, the at least one NN comprises a DQN algorithm, a DDQN algorithm, a PPO algorithm, and a SAC algorithm. By using these reinforcement training algorithms, it is possible to train the at least one NN more efficiently, which may subsequently lead to more accurate prediction of scheduling decisions for UEs.

[0023] In one example embodiment of the second aspect, the at least one NN comprises an input layer of neurons comprising a number of neurons defined by a user-defined maximum number of candidate UEs. In this embodiment, the at least one NN is configured, if a current number of candidate UEs is less than the user-defined maximum number of candidate UEs, to pad the input data with one or more zeros such that the input data has a dimension equal to the number of neurons. Since the NN topology is fixed based on the user-defined maximum number of candidate UEs, all leftover NN inputs may be padded with one or more zeros as the number of candidate UEs on the second list of UEs becomes less the user- defined maximum number of candidate UEs during said processing. Thus, also the NN outputs corresponding to the zeroed inputs may be left out from the scheduling-decision making process. For example, if the NN provides probabilities (e.g., Q-values) for each candidate UE, values of invalidated (i.e., for which RB or RBG scheduling has already been performed or is not required at all) UEs on the list of UEs may be set to minimum to avoid their selection. This may allow considering only those UEs which are be needed to be scheduled for the PRBs or RBGs, thereby reducing computational costs even more.

[0024] In one example embodiment of the second aspect, the at least one NN comprises a first NN and a second NN. The first NN is configured to obtain the scheduling decision for each of the at least one candidate UE on the second list of UEs with respect to each of said one or more RBs or RBGs. The second NN is configured to obtain a UE pairing decision for each of the at least one candidate UE on the second list of UEs with respect to each of said one or more RBs or RBGs. The UE pairing decision indicates whetherthe candidate UE needs to share said one or more RBs or RBGs with the at least one scheduled UE on the first list of UEs. With this approach, it is possible to train the first NN to optimize frequency-selective scheduling and, for example, proportional fairness, and the second NN to benefit from the MU -Ml MO aspects without compromising the performance of the UEs scheduled to the same frequency resource.

[0025] In one example embodiment of the second aspect, the input data for each of the at least one candidate UE on the second list of UEs further comprises frequency-specific information. The frequency-specific information comprises at least one of: wideband CSI for the candidate UE, sub-band CSI for the candidate UE, a number of already scheduled RBs or RBGs for the candidate UE, a past average DL throughput for the candidate UE, and buffer status information for the candidate UE. By using this additional frequency-specific information, the NN(s) may predict the scheduling decisions for UEs more accurately.

[0026] In one example embodiment of the second aspect, the at least one NN is trained based on a sequence of tuples by using a reinforcement learning algorithm. Each tuple of the sequence of tuples comprises: the input data for at least one previous candidate UE on the second list of UEs, the scheduling decision outputted by the at least one NN for each of the at least one previous candidate UE, a performance reward resulted from the scheduling decision for each of the at least one previous candidate UE, and the input data for at least one next candidate UE on the second list of UEs. By doing so, it is possible to train the NN(s) for frequency- and spatial-domain RB scheduling in the most efficient way (i.e., with good performance vs complexity characteristics).

[0027] According to a third aspect, a computer program product is provided. The computer program product comprises a computer-readable storage medium that stores a computer code. Being executed by at least one processor, the computer code causes the at least one processor to perform the method according to the second aspect. By using such a computer program product, it is possible to simplify the implementation of the method according to the second aspect in any computing device, like the apparatus according to the first aspect. According to a fourth aspect, an RB scheduling apparatus in a MU -Ml MO network is provided. The apparatus comprises a means for receiving initial information comprising: (i) a first list of UEs comprising at least one scheduled UE for which one or more RBs or RBGs have been already scheduled for a target TTI, (ii) a second list of UEs comprising at least one candidate UE for which said one or more RBs or RBGs are to be scheduled for the target TTI, (iii) a precoder for each of the at least one scheduled UE and the at least one candidate UE, and (iv) a rank for each of the at least one scheduled UE and the at least one candidate UE. The rank indicates a UE-associated number of MU-MIMO layers to be used in the MU-MIMO network. The apparatus further comprises one or more means for performing the following operations for each of the at least one candidate UE: obtaining a combined set of precoders by combining the precoder for the candidate UE with the precoder for each of the at least one scheduled UE; obtaining an array of singular values by applying SVD to the combined set of precoders; calculating a sum of ranks by summing the rank for the candidate UE with the rank for each of the at least one scheduled UE; based on the sum of ranks, selecting one or more target singular values from the array of singular values; and obtaining a scheduling decision for the candidate UE with respect to each of said one or more RBs or RBGs for the target TTI by using an ML model. The ML model comprises at least one NN configured to receive the input data comprising: (i) the sum of ranks, and (ii) said one or more target singular values, or a value derived from said one or more target singular values, and output the scheduling decision with respect to each of said one or more RBs or RBGs. The apparatus thus configured may take the MU-MIMO aspects (i.e., the spatial diversity of UEs) as well as the different ranks of the UEs in scheduling (both DL and UL) RBs or RBGs for the UEs. Furthermore, the apparatus thus configured may perform the frequency and spatial domain RB scheduling with low computational complexity regardless of: a maximum number of UEs to be scheduled to RBs or RBGs, a physical layer numerology / bandwidth (i.e., a number of RBs or RBGs per TTI), (digita l / a na log / hybrid) conventional and / or ML-based beamformers used in the MU-MIMO network, and a number of RF chains determining a maximum number of spatially scheduled UEs.

[0028] Other features and advantages of the present disclosure will be apparent upon reading the following detailed description and reviewing the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present disclosure is explained below with reference to the accompanying drawings in which:

[0030] FIG. 1 shows a block diagram of an RB scheduling apparatus in a wireless communication network in accordance with one example embodiment;

[0031] FIG. 2 shows a flowchart of a method for operating the apparatus of FIG. 2 in accordance with one example embodiment;

[0032] FIG. 3 shows one example of SVD-based value calculation and selection for each nonscheduled UE for MU -Ml MO layers;

[0033] FIG. 4 shows an exemplary architecture of a NN that may be used in the method of FIG. 2 to make RB scheduling decisions;

[0034] FIG. 5 shows a flowchart of a method for training the NN of FIG. 4 in accordance with one example embodiment;

[0035] FIG. 6 shows a flowchart of an expert-based training method for the apparatus of FIG. 1 in accordance with one example embodiment;

[0036] FIG. 7 schematically explains how two NNs may be used together in the method of FIG. 2 to make RB scheduling and UE pairing decisions;

[0037] FIG. 8 shows a Cumulative Distribution Function (CDF) versus a DL user throughput with rank adaptation and regularized zero forcing precoding MU-MIMO for different numbers of MU- MIMO layers and different RB schedulers, including the apparatus of FIG. 1; and

[0038] FIGs. 9A and 9B show a CDF versus scheduling execution time, as obtained for the apparatus of FIG. 1 and the existing RB schedulers.

[0039] DETAILED DESCRIPTION

[0040] Various embodiments of the present disclosure are further described in more detail with reference to the accompanying drawings. However, the present disclosure can be embodied in many other forms and should not be construed as limited to any certain structure or function discussed in the following description. In contrast, these embodiments are provided to make the description of the present disclosure detailed and complete.

[0041] According to the detailed description, it will be apparent to the ones skilled in the art that the scope of the present disclosure encompasses any embodiment thereof, which is disclosed herein, irrespective of whether this embodiment is implemented independently or in concert with any other embodiment of the present disclosure. For example, the apparatus and method disclosed herein can be implemented in practice by using any numbers of the embodiments provided herein. Furthermore, it should be understood that any embodiment of the present disclosure can be implemented using one or more of the elements presented in the appended claims.

[0042] Unless otherwise stated, any embodiment recited herein as "example embodiment" should not be construed as preferable or having an advantage over other embodiments.

[0043] According to the example embodiments disclosed herein, a User Equipment (UE) may refer to an electronic computing device that is configured to perform wireless communications. The UE may be implemented as a mobile station, a mobile terminal, a mobile subscriber unit, a mobile phone, a cellular phone, a smart phone, a cordless phone, a personal digital assistant (PDA), a wireless communication device, a desktop computer, a laptop computer, a tablet computer, a gaming device, a netbook, a smartbook, an ultrabook, a medical mobile device or equipment, a biometric sensor, a wearable device (e.g., a smart watch, smart glasses, a smart wrist band, etc.), an entertainment device (e.g., an audio player, a video player, etc.), a vehicular component or sensor (e.g., a driver-assistance system), a smart meter / sensor, an unmanned vehicle (e.g., an industrial robot, a quadcopter, etc.) and its component (e.g., a self-driving car computer), industrial manufacturing equipment, a global positioning system (GPS) device, an Internet-of-Things (loT) device, an Industrial loT (I loT) device, a machine-type communication (MTC) device, a group of Massive loT (MIoT) or Massive MTC (mMTC) devices / sensors, or any other suitable mobile device configured to support wireless communications. In some embodiments, the UE may refer to at least two collocated and inter-connected UEs thus defined.

[0044] As used in the example embodiments disclosed herein, a network node may refer to a node in any of a Radio Access Network (RAN) and a Core Network (CN). It should be noted that the CN may refer to a network intended for connecting different RAN nodes by providing proper interfaces therebetween. The CN may also provide a gateway to other networks, for example, a Data Network (DN).

[0045] Being part of the RAN, the network node may be implemented as a fixed point of communication / communication node for a UE in a particular wireless communication network. More specifically, the RAN node may be used to connect the UE to the DN through the CN and may be referred to as a base transceiver station (BTS) in terms of the 2G communication technology, a NodeB in terms of the 3G communication technology, an evolved NodeB (eNodeB) in terms of the 4G communication technology, and a gNB in terms of the 5G New Radio (NR) communication technology. The RAN node may serve different cells, such as a macrocell, a microcell, a picocell, a femtocell, and / or other types of cells. The macrocell may cover a relatively large geographic area (for example, at least several kilometers in radius). The microcell may cover a geographic area less than two kilometers in radius, for example. The picocell may cover a relatively small geographic area, such, for example, as offices, shopping malls, train stations, stock exchanges, etc. The femtocell may cover an even smaller geographic area (for example, a home).

[0046] Being part of the CN, the network node may also refer to any of CN network functions, such as an Access and Mobility Management Function (AMF), a Session Management Function (SMF), Unified Data Management (UDM), User Plane Function (UPF), Policy Control Function (PCF), etc. The AMF supports termination of Non-Access Stratum (NAS) signalling, NAS ciphering and integrity protection, registration management, connection management, mobility management, access authentication and authorization, security context management. The SMF supports session management (session establishment, modification, release), UE IP address allocation and management, Dynamic Host Configuration Protocol (DHCP) functions, termination of NAS signalling related to the session management, downlink (DL) data notification, traffic steering configuration for the UPF for proper traffic routing. UDM supports Authentication and Key Agreement (AKA) credentials generation, user identification handling, access authorization, subscription management. The UPF supports packet routing and forwarding, packet inspection, Quality of Service (QoS) handling, acts as an external Protocol Data Unit (PDU) session point of interconnect to the DN, and is an anchor point for intra- and inter- Radio Access Technology (RAT) mobility. The PCF supports a unified policy framework, providing policy rules to Control Plane (CP) functions, access subscription information for policy decisions in a Unified Data Repository (UDR).

[0047] In the example embodiments disclosed herein, Multi-User Multiple-Input Multiple-Output (MU-MIMO) may refer to a communication technology at which a plurality of UEs communicate with a network node by using the same time-frequency resources (i.e., resource blocks (RBs) or RB Groups (RBGs), and the network node considers communication signals from the UEs as MIMO signals and separates the signals accordingly. The MU-MIMO technology may be considered as Space-Division Multiple Access (SDMA) using a spatial channel as a resource in addition to the conventional time-frequency resources. In the MU- MIMO technology, two or more UEs are appropriately (by specially designed DL and UL schedulers usually resided in a RAN node) selected to transmit and receive data at the same time by using the same RBs or RBGs. With this, a large multi-user diversity effect can be obtained, and the cell capacity of the whole wireless communication system can be improved.

[0048] In MU-MIMO, rank adaption refers to the dynamic control of ranks of UEs according to changing channel conditions. The channel conditions may be determined by such parameters as Signal-to-lnterference-and-Noise Ratio (SINR) and fading correlation between antennae in a MU-MIMO system. With the use of spatial multiplexing, a network node (e.g., gNB) may send multiple data streams or MU-MIMO layers to the UEs in a DL transmission using the same RB or RBG. The number of such MU-MIMO layers or data streams is defined as the rank. The UE may periodically measure a channel and send a recommendation of a desired rank to the network node by using the so-called Rank Indicator (Rl). The Rl may be sent periodically or aperiodically in different schemes. Because the Rl reported to the network node may change with time, the network node may adjust the number of MU- MIMO layers used in a DL transmission forthe UE, based upon the changing Rl received from the UE. At the same time, the network node is not obliged to use the rank reported by the UE and may decide to change it, if required (e.g., based on dynamically changing network conditions).

[0049] The example embodiments disclosed herein provide a technical solution that enables RB or

[0050] RBG scheduling for UEs in a MU-MIMO network based on the spatial diversity and ranks of the UEs. For this purpose, one or more scheduled UEs for which one or more RBs or RBGs have already been scheduled for a target TTI and one or more candidate UEs for which RB or RBG scheduling has not yet been performed for the target TTI are identified. Further, for each candidate UE, a combined set of precoders is obtained by combining precoders for that candidate UE and each scheduled UE. The combined set of precoders is subjected to SVD to obtain an array of singular values, from which one or more target singular values are selected based on the sum of ranks of the candidate UE and each scheduled UE. A scheduling decision for the candidate UE is then made by using an ML model. The ML model comprises at least one NN configured to receive the sum of ranks and the target singular value(s) (or a value derived from the singular value(s)) as input data and output the scheduling decision indicating whether to schedule the RB(s) or RBG(s) to the candidate UE for the target TTI.

[0051] FIG. 1 shows a block diagram of an RB scheduling apparatus 100 in a MU-MIMO network in accordance with one example embodiment. It should be noted that the apparatus 100 may be implemented either as part of a network node (RAN or CN node) or as an individual apparatus connected to the network node by means of wire or wirelessly. As shown in FIG. 1, the apparatus 100 comprises a processor 102 and a memory 104. The memory 104 stores processor-executable instructions 106 which, when executed by the processor 102, cause the processor 102 to perform the aspects of the present disclosure, as will be described below in more detail. It should be noted that the number, arrangement, and interconnection of the constructive elements constituting the apparatus 100, which are shown in FIG. 1, are not intended to be any limitation of the present disclosure, but merely used to provide a general idea of how the constructive elements may be implemented within the apparatus 100. For example, the processor 102 may be replaced with several processors, as well as the memory 104 may be replaced with several removable and / or fixed storage devices, depending on particular applications. Furthermore, in some embodiments, the processor 102 may perform different operations required to perform data reception and transmission, such, for example, as signal modulation / demodulation, encoding / decoding, etc. Alternatively, the apparatus 100 may further comprise an individual transceiver which can be configured to perform the required operations for data reception and transmission based on commands from the processor 102. The processor 102 may be implemented as a CPU, general-purpose processor, singlepurpose processor, microcontroller, microprocessor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), digital signal processor (DSP), complex programmable logic device, etc. It should be also noted that the processor 102 may be implemented as any combination of one or more of the aforesaid. As an example, the processor 102 may be a combination of two or more microprocessors.

[0052] The memory 104 may be implemented as a classical nonvolatile or volatile memory used in the modern electronic computing machines. As an example, the nonvolatile memory may include Read-Only Memory (ROM), ferroelectric Random-Access Memory (RAM), Programmable ROM (PROM), Electrically Erasable PROM (EEPROM), solid state drive (SSD), flash memory, magnetic disk storage (such as hard drives and magnetic tapes), optical disc storage (such as CD, DVD and Blu-ray discs), etc. As for the volatile memory, examples thereof include Dynamic RAM, Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Static RAM, etc.

[0053] The processor-executable instructions 106 stored in the memory 104 may be configured as a computer-executable program code which causes the processor 102 to perform the aspects of the present disclosure. The computer-executable program code for carrying out operations or steps for the aspects of the present disclosure may be written in any combination of one or more programming languages, such as Java, C++, Python, or the like. In some examples, the computer-executable program code may be in the form of a high- level language or in a pre-compiled form and be generated by an interpreter (also pre-stored in the memory 104) on the fly.

[0054] FIG. 2 shows a flowchart of a method 200 for operating the apparatus 100 in accordance with one example embodiment.

[0055] The method 200 starts with a step S302, in which the processor 202 receives initial information comprising a first list of scheduled UEs, a second list of candidate UEs, their ranks and precoders (also referred to as precoding matrices or precoder-matrices in the art) in the MU-MIMO network. As used herein, the term "scheduled UE" refers to a UE for which one or more RBs or RBGs of a set of RBs or RBGs have already been scheduled for a target TTI by using any of the existing RB scheduling approaches, while the term "candidate UE" refers to a UE for which no RB or RBG of the set of RBs or RBGs has yet been scheduled for the target TTI. It should also be noted that the set of RBs or RBGs is shared by a set of MU -Ml MO layers supported in the MU-MIMO network, and each UE of the first and second lists of UEs may use one or more of the MU-MIMO layers for DL and UL transmissions. For this purpose, each UE may report a desired rank (e.g., by using an Rl) to a (serving) network node, and the network node may either accept it or decide to change it, if required (e.g., depending on current network conditions).

[0056] Next, the method 200 proceeds to next steps S204-S212 performed for each candidate UE on the second list of UEs in series or in parallel, depending on the capabilities of the processor 102.

[0057] In the step S204, the processor 102 obtains a combined set of precoders by combining the precoder for the candidate UE with the precoder for each scheduled UE. Further, in the step S206, the processor 102 obtains an array of singular values by applying SVD to the combined set of precoders. In the next step S208, the processor 102 sum the rank for the candidate UE with the rank for each scheduled UE. The processor 102 uses the calculated sum of ranks to select one or more target singular values from the array of singular values in the step S210. In the last step S212, the processor 102 obtains a scheduling decision for the candidate UE by using a ML model. The ML model may comprise one or more NNs configured to receive the input data comprising the sum of ranks and the target singular value(s) (and optionally, the combined set of precoders) and output the scheduling decision. In a preferred embodiment, the NN used in the method 200 is trained using any of a DQN, DDQN, PPO and SAC algorithms, or their combination. As for the scheduling decision predicted by the NN(s), it indicates whether to schedule one or more RBs or RBGs of the set of RBs or RBGs to the candidate UE for the target TTI.

[0058] Thus, by using such input data, the NN(s) may account for the spatial correction or diversity of the UEs and their different ranks to predict the scheduling decisions. In other words, the NN(s) having such input data may provide multi-rank adaption and support in the MU-MIMO network.

[0059] Preferably, the processor 102 selects a single singular value in the step S210 to keep the input data as well as the whole NN topology reasonably sized. A reasonable size helps reducing computational complexity of matrix operations that are needed to do forward passes within a real-time computing environment. At the same time, instead of one target singular value, the processor 102 may select all or subarray of singular values from the array of singular values in the step S210 based on the sum of ranks. Furthermore, either the processor 102 or the NN(s) itself may be configured to additionally perform some mathematical operation(s) on the selected target singular value(s), thereby deriving a value from the target singular value(s). This derived value may be further used in the NN input data instead of the target singular value(s). For example, the value derived from the target singular value(s) may be a sum or average of the target singular values (if at least two target singular values are selected in the step S210).

[0060] It should also be noted that the target singular value selected as input values for the NN may be used for tuning the fairness of RB or RBG scheduling. For example, a threshold value may be used for providing a zero input value to the NN if the target singular value selected in the step S210 is below the threshold value. This way the NN may provide outputs that favor cell edge performance.

[0061] FIG. 3 shows one example of the above-mentioned SVD-based value calculation and selection for each candidate UE for the MU -Ml MO layers. In this example, after the combined set of precoders obtained in the step S204 is subjected to the SVD in the step S206, an array 300 of singular values is assumed to be obtained, which is further used in the step S208 as follows. If the second list of UEs comprises one candidate UE with rank 1 and the first list of UEs comprises two scheduled UEs with ranks 2, and 1, then the sum of the ranks of the candidate and scheduled UEs is 4. Thus, in the step S208, the processor 102 may select the 4-th singular value (i.e., 0.667799) from the array 300 of singular values. If this singular value becomes too small, it means that MU-MIMO gains are suffering, and the ML model used in the method 200 needs to learn to take this into account. Instead of the single 4-th singular value, the processor 102 may also select a subarray of singular values, which includes the first to fourth singular values (i.e., 1.26084, 1.00923, 0.972511, and 0.667799), for example.

[0062] In some embodiments, the NN used in the method 200 may additionally supplement the input data for each candidate UE with one or more of the following frequency-specific information: wideband CSI for the candidate UE, sub-band CSI for the candidate UE, a number of already scheduled RBs or RBGs for the candidate UE, a past average DL throughput for the candidate UE, and buffer status information for the candidate UE. The CSI may comprise at least one of a Radio Resource Management (RRM) measurement for the RB or RBG, a RRM measurement for a sub-band comprising the RB or RBG, and a RRM measurement averaged over an entire bandwidth comprising the sub-band. The RRM measurement may comprise a SINR, a Reference Signal Received Quality (RSRQ), a Reference Signals Received Power (RSRP), a Channel Quality Indicator (CQI), and / or any other signal quality parameter. The buffer status information may be based on a transmission buffer status (reported by the candidate UE for a UL; for a DL it is known to the network node), as well as packet sizes previously transmitted and / or received.

[0063] FIG. 4 shows an exemplary architecture of a NN 400 that may be used in the method 200 to make RB scheduling decisions. The NN 400 comprises an input layer of neurons, one or more hidden layers of neurons, and an output layer of neurons. As shown in FIG. 4, there are K input values per candidate (i.e., non-scheduled) UE. One of these K values for MU-MIMO pairing is proposed to be an SVD-based input (i.e., the target singular value(s) selected in the step S210 of the method 200). More specifically, for the input layer of neurons, the K input values are provided for each candidate UE n < Nc, where Ncis the maximum number of candidate UEs considered by the NN 400. It should be noted that Ncmay be a user-defined parameter.

[0064] In one embodiment, the input layer of neurons in the NN 400 may comprise a number of neurons defined by Nc, i.e., the user-defined maximum number of candidate UEs. In this embodiment, the NN 400 may be further configured, if a current number of candidate UEs on the second list of UEs is less than the user-defined maximum number of candidate UEs (i.e., after the steps S204-S212 have already been performed for one or more candidate UEs on the second list of UEs), to pad the input data for a given candidate UE with one or more zeros such that the input data has a dimension equal to the number of neurons.

[0065] The NN 400 may be trained based on a sequence of tuples by using a reinforcement learning algorithm. Each tuple of the sequence of tuples comprises: the input data for one or more previous candidate UE on the second list of UEs, the scheduling decision outputted by the at least one NN for each previous candidate UE, a performance reward resulted from the scheduling decision for each previous candidate UE, and the input data for one or more next candidate UE on the second list of UEs.

[0066] FIG. 5 shows a flowchart of a method 500 for training the NN 400 in accordance with one example embodiment. In the method 500, it is assumed that the NN 400 is trained using the DQN or DDQN algorithm. The method 500 starts with a step S502, in which the processor 102 selects a next candidate UE on the second list of UEs. For example, said selection may be made in accordance with the order of candidate UEs on the second list of UEs, or the order of increasing / decreasing their ranks. In a next step S504, the processor 102 checks whether the selected candidate UE is valid, i.e., a non-scheduled UE. If not, the method 500 goes back to the step S502; otherwise, the method 500 goes to a step S506, in which the processor 102 adds the input data for the valid UE to an input data vector. As noted earlier, the input data includes the sum of ranks, one or more target singular values (or any value derived therefrom) and, optionally, the frequency-specific information. In a next step S508, the processor 102 checks whether the input data vector is full, i.e., whether all valid UEs are found on the second list of UEs. If not, the steps S502-S506 are repeated; otherwise, the method 500 proceeds to a step S510, in which the processor 102 generates a mask vector for the NN 400 based on the data input vector. The mask vector comprises a sequence of tuples for each valid UE, with each tuple comprising the input data for the previous valid UE on the second list of UEs, the scheduling decision outputted by the NN 400 for the previous valid UE, the performance reward resulted from the scheduling decision for the previous valid UE, and the input data for the next valid UE on the second list of UEs. In a preferred embodiment, each tuple may additionally comprise action masks for the previous and next valid UEs. For example, 0 (invalid) or 1 (valid) binary masks may be used for each possible action. In a next step S512, the processor 102 checks whether at least one other valid action is not masked out; in case of "no", the method 500 goes to a step S514, in which the processor 102 selects empty allocation, and in case of "yes", the method 500 goes to a step S514, in which the processor 102 passes the mask vector through the NN 400. In a next step S518, the processor 102 uses the output data of the NN 400 to select the best action that is not masked out.

[0067] One important outcome of the method 500 is that if all the candidate UEs are masked out and the empty allocation is the only option, then such mask vectors are obsoleted from training. Hence, the above-mentioned tuples are not formed if there is nothing to learn, i.e., the action space is squeezed already to a single possible action, which is the empty allocation. Then, the empty allocation can be done without inferring with the NN 400.

[0068] It should be noted that some of heuristic RB schedulers (especially the DL ones) perform so exhaustive calculations in a simulator environment that they are hard to beat in spectral efficiency with ML. This is because they estimate throughputs accurately by recalculating precoders, reselecting Modulation and Coding Schemes (MCSs), and recalculating throughput estimates to ensure that the scheduling decision does not decrease performance. However, this kind of exhaustiveness has a big cost in computational complexity, which makes such RB schedulers difficult or even impossible to be run in real time computing environment such as real base station products.

[0069] To tackle the problem of extensive computational complexity of the best expert RB schedulers, they may be trained into a NN (like the NN 400). Once the NN is trained to mimic a chosen expert scheduler, similar performance can be observed with quite impressive computational complexity reduction.

[0070] FIG. 6 shows a flowchart of an expert-based training method 600 for the apparatus 100 in accordance with one example embodiment. In the training method 600, a non-real time expert trainer may be implemented in each network node, or it may serve multiple network nodes. When multiple network nodes use the same non-real time expert trainer, the training method 600 may be accelerated, and the resulting accuracy may be increased. To save energy, the non-real time expert trainer may be run only when training is needed, i.e., to learn initial ML models, or occasionally to ensure that the ML models are experiencing inputs and outputs as widely as possible. The training method 600 goes as follows:

[0071] Steps S602 an S604: A real-time scheduler starts RB scheduling for an upcoming TTI and obtains inputs required forthe RB scheduling (e.g., the number of candidate UEs, the number of scheduled UEs, UE precoders and ranks, etc.).

[0072] Step S606: The real-time scheduler provides inputs that it uses for the RB scheduling to the non-real time expert trainer.

[0073] Steps S608 and S610: The non-real time expert trainer makes scheduling decisions with a deep scheduler (which is assumed to be implemented as the apparatus 100) as well as with a given expert (e.g., heuristic) scheduler. It should be noted that the deep scheduler may not use the same inputs as the expert scheduler. For example, to minimize computational complexity, only a limited number of inputs are used. Also, the inputs may be normalized to be better match the NN(s) involved in the deep scheduler.

[0074] Step S612: The decisions of the expert scheduler and the deep scheduler are compared.

[0075] Step S614: If the deep scheduler provides the same decision, it will be rewarded with a positive value (e.g., r = 1.0). If the deep scheduler makes a different decision, a reward may be zero or negative (e.g., r = -0.2).

[0076] Step S616: After that, tuples can be formed and stored to a replay memory. If the above- mentioned masking approach is used, then also the action masks for the previous and next candidate UEs are stored into the replay memory.

[0077] Step S618: The deep scheduler may be then trained with the tuples in the replay memory. Training may be performed continuously, occasionally with a desired periodicity, or when needed.

[0078] Step S620: Trained updated ML models (i.e., the weight matrices of the NN(s)) for the deep scheduler are then provided to the real-time scheduler.

[0079] Steps S622-S626: The real-time scheduler obtains scheduling decisions for the upcoming TTI together with the deep scheduler.

[0080] Step S628: The real-time scheduler allocates one or more RBs in accordance with the scheduling decisions obtained in the steps S622-S626.

[0081] This training method 600 allows training the initial deep scheduler (i.e., the apparatus 100) by using a certain expert scheduler. Further training may be done without the expert scheduler by rewarding with spectral efficiency, fairness and / or any other desired performance metric(s).

[0082] FIG. 7 schematically explains how two NNs may be used together in the method of FIG. 2 to make RB scheduling and UE pairing decisions. The first NN is trained to optimize frequency selective scheduling of a first MU-MIMO layer, and the second NN is trained to optimize UE pairing for all remaining MU-MIMO layers. The first and second NNs may be trained by individual network nodes. To speed-up their training, save costs and energy in training, and achieve more accurate scheduling decisions, multiple network nodes may be harnessed to train a single centralized ML model. This can be done by either:

[0083] Option 1: Providing pre-processed training tuples (like the ones discussed above) from all participating network nodes to a centralized entity that trains the deep scheduler and provides trained updates for each network node. This option requires the same expert scheduler running with each network node; or

[0084] Option 2: Each network node provides un-preprocessed input values used for RB scheduling, which they experience in a real environment, to a centralized entity. The centralized entity is then configured to run a chosen exhaustive expert scheduler and the deep scheduler side by side and train the deep scheduler to match the chosen expert scheduler. As in the first option, the centralized entity provides updated ML models to each network node.

[0085] One can assume that the second option would be more computational efficient at a larger scale, because the execution of the exhaustive expert scheduler is also centralized. Nevertheless, in the two training options (centralized and non-centralized) real-time computations are performed by network nodes and only direct passes of the Deep Scheduler NN may be reduced to a few fast matrix operations to minimize computational complexity (as opposed to complex heuristic algorithms that compute physics-based estimates with many recursive loops).

[0086] FIG. 8 shows a CDF versus a DL user throughput with rank adaptation and regularized zero forcing precoding MU-MIMO for different numbers of MU-MIMO layers and different RB schedulers, including the apparatus 100 (which is denoted as "Deep Scheduler"). To obtain these proof-of-concept curves, a realistic system level simulation with 21 network nodes in a 3GPP dense urban macro scenario was used with 210 UEs. The best heuristic schedulers available were used for comparison, namely:

[0087] Frequency domain (FD): Proportional Fair (PF) in ByRB mode. This scheduler loops through RBs and ensures that spectral efficiency is maximized with proportional fairness.

[0088] Spatial domain (SD): Greedy and Virtual UE. Greedy as well as Virtual UE ensures that every MU-MIMO layer for each RB increases a sum throughput for resource block. The main difference is that Virtual UE loops through RBs and pairs UEs RB by RB, while Greedy goes through RBs for one MU -Ml MO layer at a time. In these simulations, the Deep Scheduler looped as Greedy, but it can be implemented to loop similarly as so-called virtual UE as well.

[0089] All the compared schedulers use only C++ implementation and the same computational resources. Hence, there are no accelerators used for NNs. Results are apples to apples comparisons in terms of performance and computational complexity. A DDQN with 2 hidden layers (128 neurons at each hidden layer) was used for model training. The derivative of sigmoid-weighted linear unit (dSiLU) was used instead as an activation function for all layers.

[0090] As shown in FIG. 8, the Deep Scheduler was for the first time able to reach the performance of the best heuristic expert schedulers for MU-MIMO with rank adaptation enabled. This was due to implementation of the method 200.

[0091] FIGs. 9A and 9B show a CDF versus scheduling execution time, as obtained for the apparatus of FIG. 1 ("Deep Scheduler") and the existing RB schedulers. To obtain these curves, the same simulation parameters and the existing RB schedulers were used as the ones discussed above with respect to the curves of FIG. 8.

[0092] As follows from FIGs. 9A and 9B, the shown performance of the Deep Scheduler was reached with greatly decreased computational complexity. If one can reach the same performance with substantially reduced computational complexity, it is more likely that RB scheduling can be executed in a real-time computing environment within TTIs. Also, this leaves more computational resources to be used elsewhere in the system or just for saving a little bit of energy. It should be noted that computational complexity of the Deep Scheduler can be even further reduced by removing all the calculations that are not really used with the Deep Scheduler from a simulator code. Also, Deep Scheduler implementation's computational complexity can be also further optimized by optimizing the NN topology and size, reducing unnecessary high value accuracy (e.g., doubles to floats in C++ code), etc.

[0093] It should be noted that each step or operation of the methods 200 or 500, or any combinations of the steps or operations, can be implemented by various means, such as hardware, firmware, and / or software. As an example, one or more of the steps or operations described above can be embodied by processor executable instructions, data structures, program modules, and other suitable data representations. Furthermore, the processorexecutable instructions which embody the steps or operations described above can be stored on a corresponding data carrier and executed by the processor 102. This data carrier can be implemented as any computer-readable storage medium configured to be readable by said at least one processor to execute the processor executable instructions. Such computer-readable storage media can include both volatile and nonvolatile media, removable and non-removable media. By way of example, and not limitation, the computer- readable media comprise media implemented in any method or technology suitable for storing information. In more detail, the practical examples of the computer-readable media include, but are not limited to information-delivery media, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile discs (DVD), holographic media or other optical disc storage, magnetic tape, magnetic cassettes, magnetic disk storage, and other magnetic storage devices.

[0094] Although the example embodiments of the present disclosure are described herein, it should be noted that any various changes and modifications could be made in the embodiments of the present disclosure, without departing from the scope of legal protection which is defined by the appended claims. In the appended claims, the word "comprising" does not exclude other elements or operations, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

Claims

CLAIMS1. A Resource Block (RB) scheduling apparatus in a Multi-User Multiple-Input Multiple- Output (MU-MIMO) network, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: receive initial information comprising: a first list of User Equipments (UEs) comprising at least one scheduled UE for which one or more RBs or RB Groups (RBGs) have been already scheduled for a target Transmission Time Interval (TTI); a second list of UEs comprising at least one candidate UE for which said one or more RBs or RBGs are to be scheduled for the target TTI; a precoder for each of the at least one scheduled UE and each of the at least one candidate UE; and a rank for each of the at least one scheduled UE and each of the at least one candidate UE, the rank indicating a UE-associated number of MU-MIMO layers to be used in the MU-MIMO network; and for each of the at least one candidate UE: obtain a combined set of precoders by combining the precoder for the candidate UE with the precoder for each of the at least one scheduled UE; obtain an array of singular values by applying Singular Value Decomposition (SVD) to the combined set of precoders; calculate a sum of ranks by summing the rank for the candidate UE with the rank for each of the at least one scheduled UE; based on the sum of ranks, select one or more target singular values from the array of singular values; and by using a Machine-Learning (ML) model, obtain a scheduling decision for the candidate UE with respect to each of said one or more RBs or RBGs for the target TTI, the ML model comprising at least one Neural Network (NN) configured to:receive input data comprising: the sum of ranks and said one or more target singular values, or the sum of ranks and a value derived from said one or more target singular values; and output the scheduling decision with respect to each of said one or more RBs or RBGs.

2. The apparatus of claim 1, wherein the apparatus is further caused, if the scheduling decision with respect to at least one RB or RBG of said one or more RBs or RBGs indicates that the at least one RB or RBG is to be scheduled to the candidate UE, to identify the candidate UE as a scheduled UE for the at least one RB or RBG.

3. The apparatus of claim 1 or 2, wherein the at least one NN is trained using one of a Deep Q-Network (DQN) algorithm, a Double DQN (DDQN) algorithm, a Proximal Policy Optimization (PPO) algorithm, and a Soft Actor Critic (SAC) algorithm.

4. The apparatus of any one of claims 1 to 3, wherein the at least one NN comprises an input layer of neurons comprising a number of neurons defined by a user-defined maximum number of candidate UEs, and wherein the at least one NN is configured, if a current number of candidate UEs on the second list of UEs is less than the user- defined maximum number of candidate UEs, to pad the input data with one or more zeros such that the input data has a dimension equal to the number of neurons.

5. The apparatus of any one of claims 1 to 4, wherein the at least one NN comprises: a first NN configured to obtain the scheduling decision for each of the at least one candidate UE on the second list of UEs with respect to each of said one or more RBs or RBGs; and a second NN configured to obtain a UE pairing decision for each of the at least one candidate UE on the second list of UEs with respect to each of said one or more RBs or RBGs, the UE pairing decision indicating whether the candidate UE needs to share said one or more RBs or RBGs with the at least one scheduled UE on the first list of UEs.

6. The apparatus of any one of claims 1 to 5, wherein the input data for each of the at least one candidate UE on the second list of UEs further comprises frequency-specific information comprising at least one of: wideband Channel State Information (CSI) for the candidate UE; sub-band CSI for candidate UE; a number of already scheduled RBs or RBGs for the candidate UE ; a past average downlink (DL) throughput for the candidate UE; and buffer status information for the candidate UE.

7. The apparatus of any one of claims 1 to 6, wherein the at least one NN is trained based on a sequence of tuples by using a reinforcement learning algorithm, each tuple of the sequence of tuples comprising: the input data for at least one previous candidate UE on the second list of UEs; the scheduling decision outputted by the at least one NN for each of the at least one previous candidate UE; a performance reward resulted from the scheduling decision for each of the at least one previous UE; and the input data for at least one next candidate UE on the second list of UEs.

8. A Resource Block (RB) scheduling method in a Multi-User Multiple-Input Multiple- Output (MU-MIMO) network, comprising: receiving initial information comprising: a first list of User Equipments (UEs) comprising at least one scheduled UE for which one or more RBs or RB Groups (RBGs) have already been scheduled for a target Transmission Time Interval (TTI); a second list of UEs comprising at least one candidate UE for which said one or more RBs or RBGs are to be scheduled for the target TTI; a precoder for each of the at least one scheduled UE and each of the at least one candidate UE; and a rank for each of the at least one scheduled UE and each of the at least one candidate UE, the rank indicating a UE-associated number of MU-MIMO layers to be used in the MU-MIMO network; and1 for each of the at least one candidate UE: obtaining a combined set of precoders by combining the precoder for the candidate UE with the precoder for each of the at least one scheduled UE; obtaining an array of singular values by applying Singular Value Decomposition (SVD) to the combined set of precoders; calculating a sum of ranks by summing the rank for the candidate UE with the rank for each of the at least one scheduled UE; based on the sum of ranks, selecting one or more target singular values from the array of singular values; and by using a Machine-Learning (ML) model, obtaining a scheduling decision for the candidate UE with respect to each of said one or more RBs or RBGs for the target TTI, the ML model comprising at least one Neural Network (NN) configured to: receive input data comprising: the sum of ranks and said one or more target singular values, or the sum of ranks and a value derived from said one or more target singular values; and output the scheduling decision with respect to each of said one or more RBs or RBGs.

9. The method of claim 8, further comprising identifying the candidate UE as a scheduled UE for at least one RB or RBG of said one or more RBs or RBGs if the scheduling decision with respect to the at least one RB or RBG indicates that the at least one RB or RBG is to be scheduled to the candidate UE.

10. The method of claim 8 or 9, wherein the at least one NN is trained using one of a Deep Q-Network (DQN) algorithm, a Double DQN (DDQN) algorithm, a Proximal Policy Optimization (PPO) algorithm, and a Soft Actor Critic (SAC) algorithm.

11. The method of any one of claims 8 to 10, wherein the at least one NN comprises an input layer of neurons comprising a number of neurons defined by a user-defined maximum number of candidate UEs, and wherein the at least one NN is configured, if a current number of candidate UEs on the second list of UEs is less than the user-defined maximum number of candidate UEs, to pad the input data with one or more zeros such that the input data has a dimension equal to the number of neurons.

12. The method of any one of claims 8 to 11, wherein the at least one NN comprises: a first NN configured to obtain the scheduling decision for each of the candidate UE on the second list of UEs with respect to each of said one or more RBs or RBGs; and a second NN configured to obtain a UE pairing decision for each of the at least one candidate UE on the second list of UEs with respect to each of said one or more RBs or RBGs, the UE pairing decision indicating whether the candidate UE needs to share said one or more RBs or RBGs with the at least one scheduled UE on the first list of UEs.

13. The method of any one of claims 8 to 12, wherein the input data for each of the at least one candidate UE on the second list of UEs further comprises frequency-specific information comprising at least one of: wideband Channel State Information (CSI) for the candidate UE; sub-band CSI for the candidate UE; a number of already scheduled RBs or RBGs for the candidate UE; a past average downlink (DL) throughput for the candidate UE; and buffer status information for the candidate UE.

14. The method of any one of claims 8 to 13, wherein the at least one NN is trained based on a sequence of tuples by using a reinforcement learning algorithm, each tuple of the sequence of tuples comprising: the input data for at least one previous candidate UE on the second list of UEs; the scheduling decision outputted by the at least one NN for each of the at least one previous candidate UE; a performance reward resulted from the scheduling decision for each of the at least one previous candidate UE; and the input data for at least one next candidate UE on the second list of UEs.

5. A computer program product comprising a computer-readable storage medium, wherein the computer-readable storage medium stores a computer code which, when executed by at least one processor, causes the at least one processor to perform the method according to any one of claims 8 to 14.

Citation Information

Patent Citations

  • User scheduling using a graph neural network

    US20230189317A1