Coexistence method of NR-U network and WiGig network in unauthorized millimeter wave band

By building an unauthorized millimeter wave heterogeneous network and a distributed deep reinforcement learning scheduling DeepDS framework, the NR-U network can quickly adapt to dynamic environment changes, realize efficient packet transmission and resource allocation, and solve the problem of resource waste and service quality in coexistence between NR-U and WiGig networks.

CN120456037APending Publication Date: 2025-08-08GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510618490.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing coexistence method of NR-U network and WiGig network in the unauthorized millimeter band cannot quickly adapt to changes in the dynamic network environment, resulting in waste of spectrum resources and the inability to meet different network service quality requirements.

Method used

An unauthorized permissionless millimeter wave heterogeneous network is constructed, and a distributed deep reinforcement learning scheduling DeepDS framework based on constraining Markov's decision-making process is designed. The deep neural network is used to estimate the Q value and select the optimal Q value using a greedy strategy. The NR-U network transmits data packets according to the optimal Q value, dynamically updates network parameters through training, and outputs packet allocation strategies to realize data transmission of NR-U and WiGig networks.

Benefits of technology

Rapidly adjust the coexistence strategy, improve network resource utilization efficiency, enhance network operation effect, and meet the service quality requirements of different networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120456037A_ABST
    Figure CN120456037A_ABST
Patent Text Reader

Abstract

The invention provides a coexistence method of an NR-U network and a WiGig network in an unauthorized millimeter wave band, and relates to the technical field of wireless communication networks. The method comprises the following steps: firstly, constructing an unauthorized license-free millimeter wave heterogeneous network, and enabling an NR-U network and a WiGig network to share a public wireless channel to transmit a data packet; a distributed deep reinforcement learning scheduling framework DeepDS based on a constrained Markov decision process is designed based on a meta reinforcement learning algorithm, a deep neural network in the distributed deep reinforcement learning scheduling DeepDS framework is utilized to estimate a Q value, an optimal Q value is selected by using a greedy strategy, and an NR-U network transmits a data packet according to the optimal Q value and obtains an award. And updating deep neural network parameters in combination with reward feedback. And finally, inputting the state of the current NR-U and WiGig network coexistence environment based on the trained deep neural network, outputting a data packet distribution strategy, and realizing data transmission of the NR-U and WiGig networks. According to the coexistence method of the NR-U network and the WiGig network in the unauthorized millimeter wave band provided by the invention, the coexistence strategy can be quickly adjusted to meet the service quality requirements of different networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication networks, and more specifically relates to a coexistence method between an NR-U network and a WiGig network in an unlicensed millimeter wave band. Background Art

[0002] The NR-U network is a 5G New Radio (NR) technology operating in unlicensed spectrum. 5G NR technology enables standardized wireless communications in unlicensed frequency bands, enabling collaboration with existing frequency bands open to the public, such as Wi-Fi and Bluetooth. The NR-U network expands the application scope of 5G technology, enabling traditional telecom operators to independently deploy and operate private networks within a limited area, meeting the requirements for high reliability, low latency, and customized service quality. The WiGig network is a wireless communication technology based on the 802.11ad / ay standards, primarily operating in the 60 GHz frequency band. WiGig networks can provide transmission rates of up to 7 Gbps, making them suitable for high-speed data transmission over short distances. WiGig networks utilize high carrier frequencies and wide bandwidths to achieve high-speed transmission rates, while employing beamforming technology to reduce interference and increase capacity.

[0003] With the development of 5G technology, the coexistence of NR-U and WiGig networks in unlicensed millimeter wave bands has become increasingly important. NR-U networks utilize unlicensed millimeter wave bands to provide greater bandwidth and address capacity bottlenecks in densely populated environments. However, due to the inherent directionality, propagation, and blocking effects of millimeter wave frequencies, the coexistence of NR-U and WiGig networks requires sophisticated management and scheduling strategies.

[0004] Deep reinforcement learning technology has been applied to address the coexistence of NR-U and WiGig networks in unlicensed millimeter wave bands. Existing approaches to addressing coexistence between NR-U and WiGig networks in unlicensed millimeter wave bands struggle to quickly adjust coexistence strategies in the face of dynamic network environment changes. This results in inadequate utilization of unlicensed band resources, resulting in wasted spectrum resources and an inability to meet diverse network quality of service requirements. Summary of the Invention

[0005] In order to solve the problem that the existing coexistence method of NR-U network and WiGig network in unlicensed millimeter wave band cannot quickly adapt to changes in dynamic network environment, the present invention provides a coexistence method of NR-U network and WiGig network in unlicensed millimeter wave band, which can quickly adjust the coexistence strategy to meet the service quality requirements of different networks.

[0006] In order to achieve the above technical effects, the technical solutions of the present invention are as follows:

[0007] S1: Building an unlicensed, unlicensed millimeter wave heterogeneous network; in the unlicensed, unlicensed millimeter wave heterogeneous network, the NR-U network and the WiGig network share a public wireless channel to transmit data packets;

[0008] S2: Based on the meta-reinforcement learning algorithm, a distributed deep reinforcement learning scheduling framework DeepDS based on constrained Markov decision process is designed. The distributed deep reinforcement learning scheduling DeepDS framework is equipped with a deep neural network;

[0009] S3: Use a deep neural network to estimate the Q value and use a greedy strategy to select the optimal Q value. The NR-U network transmits the data packet according to the optimal Q value and obtains a reward.

[0010] S4: Train the deep neural network. During the training process, update the deep neural network parameters based on the reward.

[0011] S5: The current state of the NR-U and WiGig network coexistence environment is used as the input of the trained deep neural network, and the NR-U network's data packet allocation strategy is output. Based on the allocation strategy, the NR-U network and WiGig network perform data transmission.

[0012] Furthermore, the WiGig network includes: a multi-antenna access point AP and a single-antenna wireless client STA;

[0013] The NR-U network includes: a multi-antenna base station gNB and a single-antenna user equipment UE deployed around the single-antenna wireless client STA; in the NR-U network, the multi-antenna base station gNB directionally inputs data packets to the single-antenna user equipment UE through beamforming technology; the data transmission direction of the multi-antenna base station gNB is divided into K directions, and the single-antenna user equipment UE in each transmission direction forms a group. The multi-antenna base station gNB is equipped with an independent panel for each transmission direction, and each independent panel schedules a group of single-antenna user equipment UE.

[0014] Furthermore, the design process of the DeepDs framework for distributed deep reinforcement learning scheduling based on constrained Markov decision processes is as follows:

[0015] The coexistence problem of NR-U and wiGig networks is transformed into a reinforcement learning problem with a constrained Markov decision process. The current state of the NR-U and wiGig network coexistence environment is obtained, and the reward for the result of data transmission is defined.

[0016] Design a deep neural network and use it to estimate the Q value. The Q value represents the reward for transmitting a data packet on a single-antenna user equipment (UE) in the NR-U network, in a current NR-U and WiGig network coexistence environment. A higher Q value indicates a more reasonable data packet allocation strategy for the single-antenna user equipment (UE) in the NR-U network.

[0017] The greedy strategy is used to select the optimal Q value. The multi-antenna base station gNB in the NR-U network transmits data packets to the single-antenna user equipment UE based on the optimal Q value and obtains rewards.

[0018] Based on these technical approaches, the DeepDS framework, a distributed deep reinforcement learning scheduling framework designed based on a meta-reinforcement learning algorithm, uses deep neural networks to estimate Q values and combines this with a greedy strategy to select the optimal Q value. This allows the NR-U network to efficiently transmit data packets and earn rewards based on the optimal Q value. This improves the data transmission efficiency of the NR-U network while enhancing the flexibility of network resource allocation to meet the quality of service requirements of different networks.

[0019] Furthermore, the time for data packet transmission in the current NR-U and WiGig network coexistence environment is divided into multiple fixed-length time slots. T is used to represent the time slot set, and t is used to represent the time index, satisfying: T = {1, 2, ···, t}, t∈T. In each time slot, the WiGig network transmits data packets through the CSMA / CA mechanism. The multi-antenna base station gNB in the NR-U network uses the distributed deep reinforcement learning scheduling DeepDS framework to select multiple single-antenna user equipment UEs to transmit data packets. Any two of the multiple single-antenna user equipment UEs do not come from the group responsible for the same panel; the multi-antenna base station gNB uses the distributed deep reinforcement learning scheduling DeepDS framework for data transmission.

[0020] Furthermore, the process of training a deep neural network is:

[0021] Initialization parameters, including: learning rate, Lagrange multiplier, exploration rate of greedy strategy, experience pool, weight parameters of deep neural network, discount factor;

[0022] Input the initial state of the current NR-U network panel into the deep neural network and estimate the Q value;

[0023] Using a greedy strategy to select the optimal Q value, the multi-antenna gNB in the NR-U network transmits data packets to the single-antenna user equipment UE based on the optimal Q value and receives a reward.

[0024] The initial state of the current NR-U network panel, the result of transmitting data packets to a single-antenna user equipment (UE) based on the optimal Q value, the reward obtained during the data packet transmission, and the state of the NR-U and WiGig network coexistence environment in the next time slot are stored in the experience pool in the form of experience.

[0025] Randomly sample small batches of data from the experience pool, update the Lagrange multiplier, calculate the total reward, minimize the loss function through gradient descent, and update the deep neural network parameters.

[0026] Furthermore, the reward is calculated as follows: if a multi-antenna gNB in the NR-U network successfully transmits a data packet to a single-antenna UE, the reward is the data packet transmission rate in the current time slot; if a multi-antenna gNB in the NR-U network fails to transmit a data packet to a single-antenna UE, a fixed penalty is applied; if the NR-U network does not transmit a data packet, neither reward nor penalty is applied.

[0027] In the NR-U network, when a multi-antenna base station gNB transmits a data packet to a single-antenna user equipment UE, path loss needs to be overcome. The path loss is modeled as:

[0028]

[0029] Where n k represents the index of the single-antenna user equipment UE served by panel k, f c represents the carrier frequency, c represents the speed of light, μ represents the path loss exponent, d nk Indicates the distance between the multi-antenna base station gNB and the single-antenna user equipment UE;

[0030] Calculate the signal-to-interference-plus-noise ratio (SINR) based on the signal power received by the single-antenna user equipment (UE). If the SINR is greater than a set SINR threshold, the data packet transmission is successful; otherwise, the data packet transmission fails.

[0031] The expression for the signal-to-interference-and-noise ratio is shown as:

[0032]

[0033] Where N0 represents the noise power spectral density, W represents the bandwidth, Indicates the interference power of the WiGig network, SINR th Indicates the SINR threshold set for a single-antenna user equipment UE;

[0034] The data transmission rate between a multi-antenna gNB and a single-antenna UE is expressed as:

[0035]

[0036] Where, Indicates the transmission rate;

[0037] The expression of the reward obtained is:

[0038]

[0039] Where, represents the reward value, β represents the fixed penalty, Indicates the data packet transmission rate of the current time slot, Indicates the result of transmitting a data packet.

[0040] Furthermore, the process of updating the Lagrange multiplier is:

[0041] Define the Lagrangian function including the constraints, the expression is:

[0042]

[0043] Where L(π,λ) represents the Lagrangian function, represents the expected value of strategy π, γ l represents the discount factor, r t+1+l represents the immediate reward obtained at time step t+1+l, λ j represents the Lagrange multiplier corresponding to the j-th constraint, represents the cost of the jth constraint at time step t+1+l, C j represents the minimum cost threshold of the j-th constraint;

[0044] The experience pool randomly samples small batches of data and updates the Lagrange multiplier, which is expressed as:

[0045]

[0046] Where α2 represents the update rate of the Lagrange multiplier, represents the sampled mini-batch data, represents the average cost of all empirical data in the sampled mini-batch data, C j represents the minimum cost threshold of the j-th constraint, represents the Lagrangian function with respect to λ j gradient;

[0047] The constraints are introduced into the total reward function using Lagrange multipliers, and the total reward function considering the constraints and immediate rewards is constructed. The expression is:

[0048]

[0049] Where, ω jrepresents the indicator factor, represents the total reward function, r t+1 represents immediate reward, C j represents the minimum cost threshold of the j-th constraint, represents the cost of the jth constraint at time step t+1+l.

[0050] Furthermore, the loss function is expressed as:

[0051]

[0052] Where, represents the target Q value, Q(s i ,a i ; θ) represents the Q value output by the deep neural network, θ represents the parameters of the deep neural network, N ε represents the random sampling of small batch data in the experience pool, s i represents the current state space, a i Indicates the current action.

[0053] Furthermore, the method further includes: using a meta-reinforcement learning algorithm to train a distributed deep reinforcement learning scheduling DeepDS framework, the process of which is:

[0054] With Each task represents an unlicensed, unlicensed mmWave heterogeneous network. The meta-reinforcement learning algorithm performs multi-task training with inner and outer loops:

[0055] The inner loop training uses the loss function to update the deep neural network parameters through the gradient descent method in the same unlicensed unlicensed millimeter wave heterogeneous network. The updated expression is:

[0056]

[0057] In the formula, α represents the learning rate, θ represents the current parameter, and θ i ′ represents the updated deep neural network parameters, Represents the gradient of the loss function with respect to the parameters;

[0058] The outer loop training, for each task, uses the inner loop training of each task to update the deep neural network parameters, calculate the meta-loss gradient, update the initial deep neural network parameters, and minimize the cross-task loss.

[0059] Furthermore, the multi-antenna gNB uses the distributed deep reinforcement learning scheduling DeepDS framework for data transmission, which is divided into three phases:

[0060] Phase 1: The multi-antenna gNB uses a deep neural network to generate a data packet allocation strategy for each panel. If the panel's allocation strategy is to transmit data packets to a single-antenna UE, the gNB adjusts its beamforming direction to align with the selected single-antenna UE for data transmission.

[0061] Phase 2: The single-antenna UE sends an acknowledgment signal to the multi-antenna gNB. If the single-antenna UE receives a data packet, it sends an ACK signal. If the single-antenna UE does not receive a data packet, it sends a NACK signal.

[0062] Phase 3: The single-antenna user equipment (UE) that receives the data packet measures the interference power of the WiGig network and other NR-U panels in the current time slot and feeds back the detected interference power to the multi-antenna base station (gNB).

[0063] Compared with the prior art, the beneficial effects of this method are:

[0064] The present invention provides a coexistence method for NR-U networks and WiGig networks in unlicensed millimeter wave bands. First, an unlicensed, unlicensed millimeter wave heterogeneous network is constructed, in which the NR-U network and the WiGig network share a common wireless channel to transmit data packets. Then, based on the meta-reinforcement learning algorithm, a distributed deep reinforcement learning scheduling DeepDS framework based on a constrained Markov decision process is designed, in which a deep neural network is embedded. The Q value is estimated by the deep neural network, and the optimal Q value is selected using a greedy strategy. The NR-U network transmits data packets based on the optimal Q value and obtains rewards. During the training of the deep neural network, the network parameters are dynamically updated according to the rewards. Finally, based on the trained deep neural network input, the state of the current NR-U and WiGig network coexistence environment is output, and the data packet allocation strategy is output to realize data transmission between the NR-U and WiGig networks. This method can quickly adjust the coexistence strategy to meet the service quality requirements of different networks, thereby improving the utilization efficiency of network resources and enhancing the overall operation effect of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 A flowchart illustrating a method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band proposed in an embodiment of the present invention;

[0066] Figure 2 A schematic diagram of an unlicensed millimeter wave heterogeneous network proposed in an embodiment of the present invention is shown;

[0067] Figure 3 A schematic diagram illustrating the DeepDs framework for distributed deep reinforcement learning scheduling proposed in an embodiment of the present invention;

[0068] Figure 4 A schematic diagram illustrating data transmission performed by a multi-antenna base station gNB proposed in an embodiment of the present invention is shown;

[0069] Figure 5 Pseudo code representing the DeepDs framework for distributed deep reinforcement learning scheduling proposed in an embodiment of the present invention;

[0070] Figure 6 The figure shows a learning flow chart of the meta-reinforcement learning algorithm proposed in an embodiment of the present invention. DETAILED DESCRIPTION

[0071] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;

[0072] In order to better illustrate this embodiment, some parts of the drawings may be omitted, enlarged, or reduced, and do not represent the actual size;

[0073] It is understandable to those skilled in the art that descriptions of certain well-known contents may be omitted in the drawings.

[0074] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0075] The positional relationships described in the drawings are for illustrative purposes only and should not be construed as limiting this patent;

[0076] Example 1

[0077] This embodiment proposes a coexistence method between NR-U network and WiGig network in unlicensed millimeter wave band, such as Figure 1 The method proposed in this embodiment generally includes the following steps:

[0078] S1: Building an unlicensed, unlicensed millimeter wave heterogeneous network; in the unlicensed, unlicensed millimeter wave heterogeneous network, the NR-U network and the WiGig network share a public wireless channel to transmit data packets;

[0079] S2: Based on the meta-reinforcement learning algorithm, a distributed deep reinforcement learning scheduling framework DeepDS based on constrained Markov decision process is designed. The distributed deep reinforcement learning scheduling DeepDS framework is equipped with a deep neural network;

[0080] S3: Use a deep neural network to estimate the Q value and use a greedy strategy to select the optimal Q value. The NR-U network transmits the data packet according to the optimal Q value and obtains a reward.

[0081] S4: Train the deep neural network. During the training process, update the deep neural network parameters based on the reward.

[0082] S5: The current state of the NR-U and WiGig network coexistence environment is used as the input of the trained deep neural network, and the NR-U network's data packet allocation strategy is used as the output. Based on the allocation strategy, the NR-U network and WiGig network perform data transmission.

[0083] In this embodiment, the WiGig network includes: a multi-antenna access point AP and a single-antenna wireless client STA; the WiGig network uses the CSMA / CA mechanism to transmit data packets;

[0084] In this embodiment, if Figure 2 The diagram shows an unlicensed, unlicensed millimeter wave heterogeneous network. The NR-U network includes: multi-antenna base stations gNB and single-antenna user equipment UE deployed around a single-antenna wireless client STA; in the NR-U network, the multi-antenna base station gNB directionally inputs data packets to the single-antenna user equipment UE through beamforming technology; the multi-antenna base station gNB data transmission direction is divided into K directions, and the single-antenna user equipment UE in each transmission direction forms a group. The multi-antenna base station gNB is equipped with an independent panel for each transmission direction, and each independent panel schedules a group of single-antenna user equipment UE.

[0085] Multi-antenna gNBs transmit data packets using beamforming technology to overcome the high path loss in the millimeter wave band. Beamforming is an antenna technology that focuses the energy of wireless signals so that they are transmitted in a specific direction, thereby increasing signal coverage and strength.

[0086] In this embodiment, the WiGig network includes M multi-antenna access points (APs). The NR-U network includes N single-antenna user equipment (UEs). The multi-antenna base station gNB has a total of K independent panels, each of which is responsible for serving a specific single-antenna user equipment (UE) group k (k∈K={1,2,····,K}). The single-antenna user equipment (UE) group k includes N UEs located in a fixed area. k single-antenna user equipment UE, and does not interfere with other panels, where

[0087] Each panel is controlled by an RF radio chain, and the multi-antenna base station gNB only activates one single-antenna user equipment UE in a group at a time through the RF beamforming vector to avoid interference between different beams.

[0088] In the NR-U network, multi-antenna gNBs transmit data packets in a directional manner using beamforming to overcome the high path loss in the mmWave frequency band. Single-antenna user equipment (UEs) receive data packets in an omnidirectional manner. Synchronization is maintained between the multi-antenna gNB and all single-antenna UEs, while synchronization between the NR-U network and the WiGig network is not required. Therefore, a time slot system is designed, where t∈T={1,2,···,t}. t represents a time index identifying a specific time slot in the time slot system, and T represents a set of time slots. In each time slot, the WiGig network transmits data packets using the CSMA / CA mechanism. The multi-antenna gNBs in the NR-U network utilize the distributed deep reinforcement learning scheduling framework, DeepDS, to select multiple single-antenna UEs for data packet transmission. No two of these multiple UEs are from the same group. Due to the directional transmission, the multi-antenna gNBs can successfully transmit data packets to selected UEs in different directions simultaneously, sharing the same unlicensed mmWave frequency band with the WiGig network.

[0089] In this embodiment, the design process of the DeepDs framework for distributed deep reinforcement learning scheduling based on constrained Markov decision processes is as follows:

[0090] The coexistence problem of NR-U and wiGig networks is transformed into a reinforcement learning problem with a constrained Markov decision process. The current state of the NR-U and wiGig network coexistence environment is obtained, and the reward for the result of data transmission is defined.

[0091] Design a deep neural network and use it to estimate the Q value. The Q value represents the reward for transmitting a data packet on a single-antenna user equipment (UE) in the NR-U network, in a current NR-U and WiGig network coexistence environment. A higher Q value indicates a more reasonable data packet allocation strategy for the single-antenna user equipment (UE) in the NR-U network.

[0092] The greedy strategy is used to select the optimal Q value. The multi-antenna base station gNB in the NR-U network transmits data packets to the single-antenna user equipment UE based on the optimal Q value and obtains rewards.

[0093] Multiple deep neural networks in the DeepDs framework, a distributed deep reinforcement learning scheduling framework, independently make decisions for different panels. This allows each panel to make scheduling decisions based on its own local information, without requiring global network information. This minimizes interference with the WiGig system while maximizing the total data rate of the NR-U network and meeting the Quality of Service (QoS) requirements of each single-antenna user equipment (UE).

[0094] In this embodiment, the antenna base station gNB acts as a reinforcement learning agent, performing parallel scheduling decisions for different panels in a separate manner through different deep neural networks DNN.

[0095] For example, Figure 3 Figure 3 is a schematic diagram of the distributed deep reinforcement learning scheduling DeepDs framework. For a specific panel k, this embodiment defines its state based on historical transmission results and detected interference power levels. The agent defines an action for each panel, i.e., selecting a single-antenna user equipment UE for data packet transmission or remaining idle. The reward function is designed based on the successful transmission, failed transmission, or idle state of the single-antenna user equipment UE. The cost function is related to the data rate requirement of the UE and is used to ensure that the quality of service requirements are met. The agent uses an ∈-greedy strategy to balance exploration (trying new data transmission) and exploitation (selecting the known best Q value for data packet transmission).

[0096] The data transmission time is divided into multiple fixed-length time slots. In each time slot, the multi-antenna gNB uses the distributed deep reinforcement learning scheduling DeepDS framework to transmit data.

[0097] like Figure 4 The following figure shows a schematic diagram of data transmission performed by a multi-antenna base station gNB. Data transmission is divided into three phases:

[0098] Phase 1: The multi-antenna gNB uses a deep neural network to generate a data packet allocation strategy for each panel. If the panel's allocation strategy is to transmit data packets to a single-antenna UE, the gNB adjusts its beamforming direction to align with the selected single-antenna UE for data transmission.

[0099] Phase 2: The single-antenna UE sends an acknowledgment signal to the multi-antenna gNB. If the single-antenna UE receives a data packet, it sends an ACK signal. If the single-antenna UE does not receive a data packet, it sends a NACK signal.

[0100] Phase 3: The single-antenna user equipment UE that receives the data packet measures the interference power of the WiGig network and other NR-U panels in the current time slot. The detected power information is given by Indicates that the detected interference power is fed back to the multi-antenna base station gNB.

[0101] The goal of a multi-antenna gNB is to maximize the total data rate of the NR-U network while minimizing interference to the WiGig system. To ensure the quality of service (QoS) of each single-antenna user equipment (UE), the long-term average data rate of each single-antenna UE must be greater than a certain value. Multi-antenna gNBs lack information about the WiGig network. Without direct knowledge of the interference level of the NR-U network on the WiGig network, a distributed deep reinforcement learning scheduling framework, DeepDs, is designed to enable coexistence between the NR-U and WiGig networks.

[0102] The present invention uses deep reinforcement learning (DRL) technology to develop a user equipment (UE) scheduling solution, namely the distributed deep reinforcement learning scheduling DeepDS framework. Figure 5 The pseudo code for the distributed deep reinforcement learning scheduling DeepDs framework shown in the figure above is used to implement the distributed deep reinforcement learning scheduling DeepDS framework. The distributed deep reinforcement learning scheduling DeepDS framework uses multiple deep neural networks (DNNs) to make decisions, optimize beacon scheduling and maximize the total data rate of the NR-U network, while reducing interference to the WiGig network and meeting the quality of service (QoS) requirements of each user equipment UE. The present invention formulates the coexistence problem as a constrained Markov decision process (CMDP) framework and proposes a new DRL algorithm that incorporates Lagrangian primal-dual optimization into the deep Q network framework, called adaptive multi-constraint deep Q network (AMC-DQN). The AMC-DQN algorithm enables the distributed deep reinforcement learning scheduling DeepDS framework to maximize the network data rate while meeting the quality of service requirements without obtaining previous operations of the WiGig network.

[0103] Example 2

[0104] In this embodiment, the process of training a deep neural network is described in detail. The process of training a deep neural network is:

[0105] Initialization parameters, including: learning rate, Lagrange multiplier, exploration rate of greedy strategy, experience pool, weight parameters of deep neural network, discount factor;

[0106] Input the initial state of the current NR-U network panel into the deep neural network and estimate the Q value;

[0107] Using a greedy strategy to select the optimal Q value, the multi-antenna gNB in the NR-U network transmits data packets to the single-antenna user equipment UE based on the optimal Q value and receives a reward.

[0108] The initial state of the current NR-U network panel, the result of transmitting data packets to a single-antenna user equipment (UE) based on the optimal Q value, the reward obtained during the data packet transmission, and the state of the NR-U and WiGig network coexistence environment in the next time slot are stored in the experience pool in the form of experience.

[0109] Randomly sample small batches of data from the experience pool, update the Lagrange multiplier, calculate the total reward, minimize the loss function through gradient descent, and update the deep neural network parameters.

[0110] Through the experience pool, the experience replay mechanism is used to store and reuse experience, which helps to stabilize the training of deep neural networks (DNNs).

[0111] The reward is calculated as follows: if a multi-antenna gNB successfully transmits a data packet to a single-antenna UE in the NR-U network, the reward is the data packet transmission rate in the current timeslot. If a multi-antenna gNB fails to transmit a data packet to a single-antenna UE in the NR-U network, a fixed penalty is applied. If no data packet is transmitted in the NR-U network, neither reward nor penalty is applied.

[0112] In the NR-U network, when a multi-antenna base station gNB transmits a data packet to a single-antenna user equipment UE, path loss needs to be overcome. The path loss is modeled as:

[0113]

[0114] Where n k represents the index of the single-antenna user equipment UE served by panel k, f c represents the carrier frequency, c represents the speed of light, μ represents the path loss exponent, d nk Indicates the distance between the multi-antenna base station gNB and the single-antenna user equipment UE;

[0115] Calculate the signal-to-interference-plus-noise ratio (SINR) based on the signal power received by the single-antenna user equipment (UE). If the SINR is greater than a set SINR threshold, the data packet transmission is successful; otherwise, the data packet transmission fails.

[0116] The expression for the signal-to-interference-and-noise ratio is shown as:

[0117]

[0118] Where N0 represents the noise power spectral density, W represents the bandwidth, Indicates the interference power of the WiGig network, SINR th Indicates the SINR threshold set for a single-antenna user equipment UE;

[0119] Interference power of WiGig networks The expression is:

[0120]

[0121] Where ξ represents the small-scale fading component, which remains constant within a single time slot but varies between different time slots. tx represents the transmit power of the multi-antenna base station gNB, G tx represents the transmit antenna gain of the multi-antenna base station gNB, Represents the receiving antenna gain of a single-antenna user equipment UE. Since the multi-antenna base station gNB adopts directional transmission, G tx According to the shape and direction of the beam is set to G m or G s .

[0122] The data transmission rate between a multi-antenna gNB and a single-antenna UE is expressed as:

[0123]

[0124] Where, Indicates the transmission rate.

[0125] The expression of the reward obtained is:

[0126]

[0127] Where, represents the reward value, β represents the fixed penalty, Indicates the data packet transmission rate of the current time slot, Indicates the result of transmitting a data packet.

[0128] In this embodiment, the NR-U network needs to minimize interference with the WiGig network. Therefore, the present invention introduces the CMDP (Constrained Markov Decision Process) framework to improve the MDP by imposing additional constraints. The multi-constraint problem in the CMDP is handled through Lagrangian primal-dual optimization. Constraints are integrated into the objective function, transforming the CMDP problem into an unconstrained problem.

[0129] In the CMDP paradigm, multi-constraint problems can be represented by a sextuple (S, A, P, R, C, γ), where S represents the state space; A represents the action space; P represents the transition probability; R represents the reward function, which determines the reward corresponding to a specific action in a given state; γ represents the discount factor; C = {c j |j∈{1,2,…,J}} represents the set of proxy cost functions that determine the cost of all constraints in each time slot, where J is the number of cost functions.

[0130] In this embodiment, the process of updating the Lagrange multiplier is:

[0131] Define the Lagrangian function including the constraints, the expression is:

[0132]

[0133] Where L(π,λ) represents the Lagrangian function, represents the expected value of strategy π, γ l represents the discount factor, r t+1+l represents the immediate reward obtained at time step t+1+l, λ j represents the Lagrange multiplier corresponding to the j-th constraint, represents the cost of the jth constraint at time step t+1+l, C j represents the minimum cost threshold of the j-th constraint;

[0134] The experience pool randomly samples small batches of data and updates the Lagrange multiplier, which is expressed as:

[0135]

[0136] Where α2 represents the update rate of the Lagrange multiplier, represents the sampled mini-batch data, represents the average cost of all empirical data in the sampled mini-batch data, C j represents the minimum cost threshold of the j-th constraint, represents the Lagrangian function with respect to λ j gradient.

[0137] The constraints are introduced into the total reward function using Lagrange multipliers, and the total reward function considering the constraints and immediate rewards is constructed. The expression is:

[0138]

[0139] Where, ω j represents the indicator factor, represents the total reward function, r t+1 represents immediate reward, C j represents the minimum cost threshold of the j-th constraint, represents the cost of the jth constraint at time step t+1+l.

[0140] The loss function is expressed as:

[0141]

[0142] Where, represents the target Q value, Q(s i ,ai ; θ) represents the Q value output by the deep neural network, θ represents the parameters of the deep neural network, N ε represents the random sampling of small batch data in the experience pool, s i represents the current state space, a i Indicates the current action.

[0143] In this embodiment, a dual-time-scale optimization strategy is adopted, in which the update speed of the deep neural network parameter θ is faster than the update speed of the Lagrange multiplier λ to ensure the stability of the algorithm.

[0144] Example 3

[0145] This embodiment provides another method for the coexistence of NR-U networks and WiGig networks in unlicensed millimeter wave bands, which, based on Example 2, also includes: using a meta-reinforcement learning algorithm to train a distributed deep reinforcement learning scheduling DeepDS framework.

[0146] The core of meta-reinforcement learning is to enable agents to master the ability to "learn how to learn," enabling them to quickly adapt to new tasks or environments. Based on the DeepDs framework for distributed deep reinforcement learning scheduling, a meta-reinforcement learning algorithm is introduced to represent the deep neural network (DNN) deployed on each panel. This algorithm not only learns to make optimal decisions in the current unlicensed, unlicensed, and heterogeneous millimeter-wave network environment, but also rapidly adjusts its strategy to address sudden changes in the unlicensed, unlicensed, and heterogeneous millimeter-wave network environment, such as shifts in the distribution of single-antenna user equipment (UE) or updates to quality of service requirements.

[0147] Meta-reinforcement learning accelerates the mastery of new tasks by learning common feature representations. These features can be quickly adapted to new environments, reducing the time required to retrain deep neural networks. At the same time, meta-reinforcement learning helps deep neural networks (DNNs) identify commonalities between different tasks, enabling knowledge transfer from one task to another.

[0148] like Figure 6 The learning flow chart of the meta-reinforcement learning algorithm shown in the figure shows the process of using the meta-reinforcement learning algorithm to train the distributed deep reinforcement learning scheduling DeepDS framework:

[0149] With Each task represents an unlicensed, unlicensed mmWave heterogeneous network. The meta-reinforcement learning algorithm performs multi-task training with inner and outer loops:

[0150] The inner loop training takes the current state of the NR-U and WiGig network coexistence environment as the input of the deep neural network, collects the current network state, the result of transmitting data packets to the single-antenna user equipment UE according to the optimal Q value, the reward obtained during the transmission of the data packet, and the state sequence of the NR-U and WiGig network coexistence environment in the next time slot (s t ,a t ,r t ,s t+1 ).

[0151] The inner loop training uses the loss function to update the deep neural network parameters through the gradient descent method in the same unlicensed unlicensed millimeter wave heterogeneous network. The updated expression is:

[0152]

[0153] In the formula, α represents the learning rate, θ represents the current parameter, and θ i ′ represents the updated deep neural network parameters, Represents the gradient of the loss function with respect to the parameters.

[0154] The outer loop training, for each task, uses the inner loop training of each task to update the deep neural network parameters, calculate the meta-loss gradient, update the initial deep neural network parameters, and minimize the cross-task loss.

[0155] This paper incorporates the concept of meta-reinforcement learning to improve the adaptability and generalization capabilities of the DeepDS framework for distributed deep reinforcement learning scheduling. In scenarios where NR-U networks coexist with WiGig networks, meta-reinforcement learning can help the DeepDS framework quickly adapt to environmental changes, such as changes in user density, blockers, and antenna array configurations, thereby dynamically adjusting the coexistence strategy.

[0156] Meta-reinforcement learning helps deep neural networks (DNNs) adapt to new changes in real time in an environment where NR-U and WiGig networks coexist. It adjusts the NR-U network's packet allocation strategy to address changes in single-antenna UE behavior, the emergence of new interference sources, or changes in network topology, thereby maintaining optimal packet allocation. This not only enables efficient single-antenna UE scheduling but also maintains robustness in a changing environment, improving the performance and efficiency of NR-U networks in unlicensed millimeter wave bands.

[0157] The embodiments are provided merely to illustrate the present invention and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications may be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the claims.

Claims

1. A method for coexistence of NR-U network and WiGig network in unlicensed millimeter wave band, characterized in that: The following steps are involved: S1: Building an unlicensed, unlicensed millimeter wave heterogeneous network; in the unlicensed, unlicensed millimeter wave heterogeneous network, the NR-U network and the WiGig network share a public wireless channel to transmit data packets; S2: Based on the meta-reinforcement learning algorithm, a distributed deep reinforcement learning scheduling framework DeepDS based on constrained Markov decision process is designed. The distributed deep reinforcement learning scheduling DeepDS framework is equipped with a deep neural network; S3: Use a deep neural network to estimate the Q value and use a greedy strategy to select the optimal Q value. The NR-U network transmits the data packet according to the optimal Q value and obtains a reward. S4: Train the deep neural network. During the training process, update the deep neural network parameters based on the reward. S5: The current state of the NR-U and WiGig network coexistence environment is used as the input of the trained deep neural network, and the NR-U network's data packet allocation strategy is output. Based on the allocation strategy, the NR-U network and WiGig network perform data transmission.

2. The method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band according to claim 1, wherein: The WiGig network includes: a multi-antenna access point AP and a single-antenna wireless client STA; The NR-U network includes: a multi-antenna base station gNB and a single-antenna user equipment UE deployed around the single-antenna wireless client STA; in the NR-U network, the multi-antenna base station gNB directionally inputs data packets to the single-antenna user equipment UE through beamforming technology; the data transmission direction of the multi-antenna base station gNB is divided into K directions, and the single-antenna user equipment UE in each transmission direction forms a group. The multi-antenna base station gNB is equipped with an independent panel for each transmission direction, and each independent panel schedules a group of single-antenna user equipment UE.

3. The method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band according to claim 2, wherein: The design process of the DeepDs framework for distributed deep reinforcement learning scheduling based on constrained Markov decision processes is as follows: The coexistence problem of NR-U and wiGig networks is transformed into a reinforcement learning problem with a constrained Markov decision process. The current state of the NR-U and wiGig network coexistence environment is obtained, and the reward for the result of data transmission is defined. Design a deep neural network and use it to estimate the Q value. The Q value represents the reward for transmitting a data packet on a single-antenna user equipment (UE) in the NR-U network, in a current NR-U and WiGig network coexistence environment. A higher Q value indicates a more reasonable data packet allocation strategy for the single-antenna user equipment (UE) in the NR-U network. The greedy strategy is used to select the optimal Q value. The multi-antenna base station gNB in the NR-U network transmits data packets to the single-antenna user equipment UE based on the optimal Q value and obtains rewards.

4. The method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band according to claim 2, wherein: The data packet transmission time in the current NR-U and WiGig network coexistence environment is divided into multiple fixed-length time slots. T is used to represent the time slot set and t is used to represent the time index, satisfying: T = {1, 2, ···, t}, t∈T. In each time slot, the WiGig network transmits data packets using the CSMA / CA mechanism. The multi-antenna gNB in the NR-U network uses the distributed deep reinforcement learning scheduling DeepDS framework to select multiple single-antenna user equipment (UE) to transmit data packets. Any two of the multiple single-antenna user equipment (UE) are not from the group managed by the same panel. The multi-antenna base station gNB uses the distributed deep reinforcement learning scheduling DeepDS framework for data transmission.

5. The method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band according to claim 4, characterized in that: The process of training a deep neural network is: Initialization parameters, including: learning rate, Lagrange multiplier, exploration rate of greedy strategy, experience pool, weight parameters of deep neural network, discount factor; Input the initial state of the current NR-U network panel into the deep neural network and estimate the Q value; Using a greedy strategy to select the optimal Q value, the multi-antenna gNB in the NR-U network transmits data packets to the single-antenna user equipment UE based on the optimal Q value and receives a reward. The initial state of the current NR-U network panel, the result of transmitting data packets to a single-antenna user equipment (UE) based on the optimal Q value, the reward obtained during the data packet transmission, and the state of the NR-U and WiGig network coexistence environment in the next time slot are stored in the experience pool in the form of experience. Randomly sample small batches of data from the experience pool, update the Lagrange multiplier, calculate the total reward, minimize the loss function through gradient descent, and update the deep neural network parameters.

6. The method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band according to claim 5, characterized in that: The reward is calculated as follows: if a multi-antenna gNB successfully transmits a data packet to a single-antenna UE in the NR-U network, the reward is the data packet transmission rate in the current timeslot. If a multi-antenna gNB fails to transmit a data packet to a single-antenna UE in the NR-U network, a fixed penalty is applied. If no data packet is transmitted in the NR-U network, neither reward nor penalty is applied. In the NR-U network, when a multi-antenna base station gNB transmits a data packet to a single-antenna user equipment UE, path loss needs to be overcome. The path loss is modeled as: Where n k represents the index of the single-antenna user equipment UE served by panel k, f c represents the carrier frequency, c represents the speed of light, μ represents the path loss exponent, d nk Indicates the distance between the multi-antenna base station gNB and the single-antenna user equipment UE; Calculate the signal-to-interference-plus-noise ratio (SINR) based on the signal power received by the single-antenna user equipment (UE). If the SINR is greater than a set SINR threshold, the data packet transmission is successful; otherwise, the data packet transmission fails. The expression for the signal-to-interference-and-noise ratio is shown as: Where N0 represents the noise power spectral density, W represents the bandwidth, Indicates the interference power of the WiGig network, SINR th Indicates the SINR threshold set for a single-antenna user equipment UE; The data transmission rate between a multi-antenna gNB and a single-antenna UE is expressed as: Where, Indicates the transmission rate; The expression of the reward obtained is: Where, represents the reward value, β represents the fixed penalty, Indicates the data packet transmission rate of the current time slot, Indicates the result of transmitting a data packet.

7. The method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band according to claim 6, characterized in that: The process of updating the Lagrange multiplier is: Define the Lagrangian function including the constraints, the expression is: Where L(π,λ) represents the Lagrangian function, represents the expected value of strategy π, γ l represents the discount factor, r t+1+l represents the immediate reward obtained at time step t+1+l, λ j represents the Lagrange multiplier corresponding to the j-th constraint, represents the cost of the jth constraint at time step t+1+l, C j represents the minimum cost threshold of the j-th constraint; The experience pool randomly samples small batches of data and updates the Lagrange multiplier, which is expressed as: Where α2 represents the update rate of the Lagrange multiplier, B represents the sampled mini-batch data, represents the average cost of all empirical data in the sampled mini-batch data, C j represents the minimum cost threshold of the j-th constraint, represents the Lagrangian function with respect to λ j gradient; The constraints are introduced into the total reward function using Lagrange multipliers, and the total reward function considering the constraints and immediate rewards is constructed. The expression is: Where, ω j represents the indicator factor, represents the total reward function, r t+1 represents immediate reward, C j represents the minimum cost threshold of the j-th constraint, represents the cost of the jth constraint at time step t+1+l.

8. The method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band according to claim 7, characterized in that: The loss function is expressed as: Where, Indicates the target Q value, Q(s i ,a i ; θ) represents the Q value output by the deep neural network, θ represents the parameters of the deep neural network, N ε represents the random sampling of small batch data in the experience pool, s i represents the current state space, a i Indicates the current action.

9. The method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band according to claim 8, characterized in that: It also includes: using the meta-reinforcement learning algorithm to train the distributed deep reinforcement learning scheduling DeepDS framework. The process is: With Each task represents an unlicensed, unlicensed mmWave heterogeneous network. The meta-reinforcement learning algorithm performs multi-task training with inner and outer loops: The inner loop training uses the loss function to update the deep neural network parameters through the gradient descent method in the same unlicensed unlicensed millimeter wave heterogeneous network. The updated expression is: In the formula, α represents the learning rate, θ represents the current parameter, and θ i ′ represents the updated deep neural network parameters, Represents the gradient of the loss function with respect to the parameters; The outer loop training, for each task, uses the inner loop training of each task to update the deep neural network parameters, calculate the meta-loss gradient, update the initial deep neural network parameters, and minimize the cross-task loss.

10. The method for coexistence of an NR-U network and a WiGig network in an unlicensed millimeter wave band according to claim 4, characterized in that: The multi-antenna gNB uses the distributed deep reinforcement learning scheduling framework DeepDS for data transmission. The data transmission is divided into three phases: Phase 1: The multi-antenna gNB uses a deep neural network to generate a data packet allocation strategy for each panel. If the panel's allocation strategy is to transmit data packets to a single-antenna UE, the gNB adjusts its beamforming direction to align with the selected single-antenna UE for data transmission. Phase 2: The single-antenna UE sends an acknowledgment signal to the multi-antenna gNB. If the single-antenna UE receives a data packet, it sends an ACK signal. If the single-antenna UE does not receive a data packet, it sends a NACK signal. Phase 3: The single-antenna user equipment (UE) that receives the data packet measures the interference power of the WiGig network and other NR-U panels in the current time slot and feeds back the detected interference power to the multi-antenna base station (gNB).