Wireless Resource Allocation Method, Model Training and Inference Methods, Devices and Equipment
The online training iterator updates the solver parameters and combines dynamic channel gain to allocate OFDMA network resources, which solves the problem of unreasonable resource allocation in the existing technology, improves network performance, and has strong environmental adaptability.
Patent Information
- Application Number
- CN202110507804.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-05-10
AI Technical Summary
The prior art is difficult to effectively allocate resources in OFDMA networks, resulting in insufficient network performance.
By obtaining the large-scale and small-scale channel gains in the dynamic environment, combining the parameters of the solver at the previous moment and the target value of the small-scale channel gain, the online training iterator updates the solver parameters and allocates resources according to the updated parameters.
The rational allocation of OFDMA network resources is achieved, the performance of the entire network is enhanced, and the neural network is simple in structure, fast real-time decision-making speed, and strong ability to adapt to environmental changes.
Smart Images

Figure CN115334533B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technologies, and in particular, to a wireless resource allocation method, a machine learning model training and inference method adapted to dynamic changes in the communication environment, a device, a network-side device, and a storage medium. Background Art
[0002] Wireless communication technologies aim to provide higher data transmission rates and diverse wireless interfaces to ensure continuous seamless connection for all users. However, the battery capacity of existing terminals lags far behind the battery capacity required by those rich multimedia services and multi-mode multi-standard services. Therefore, energy-efficient transmission, also widely known as green communication, has attracted more and more attention.
[0003] Orthogonal Frequency Division Multiple Access (OFDMA) technology is one of the mainstream multiple access schemes adopted by future broadband wireless networks, such as The Third Generation Partnership Project (3GPP), Long Term Evolution (LTE), etc. However, how to reasonably allocate OFDMA network resources to enhance the performance of the entire network has become an urgent problem to be solved. Summary of the Invention
[0004] The wireless resource allocation method, the machine learning model training and inference method adapted to dynamic changes in the communication environment, the device, the network-side device, and the storage medium provided by the present application can reasonably allocate OFDMA network resources to different users, thereby enhancing the performance of the entire network.
[0005] According to a first aspect of the present application, there is provided a wireless resource allocation method, including:
[0006] Obtain the large-scale channel gain at the current moment in the dynamic environment and the small-scale channel gain at the current moment;
[0007] Obtain the target value of the small-scale channel gain, and obtain the parameters of the solver at the previous moment; the parameters include user bandwidth, multiplier, and power neural network parameters;
[0008] Online train an iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment;
[0009] Perform resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment.
[0010] In some embodiments of the present application, obtaining the target value of the small-scale channel gain includes:
[0011] Obtaining the small-scale channel gain from the offline environment, and using the small-scale channel gain obtained from the offline environment as the target value of the small-scale channel gain.
[0012] In some embodiments of the present application, online training the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment includes:
[0013] Calculating the corresponding gradient according to the large-scale channel gain at the current moment;
[0014] Using the parameters of the solver at the previous moment as the initial value of the iterator at the current moment;
[0015] Performing N iterations on the iterator at the current moment according to the corresponding gradient and the target value of the small-scale channel gain;
[0016] Using the parameters obtained after the N iterations of the iterator at the current moment as the parameters of the solver at the current moment.
[0017] In an embodiment of the present application, the formula of the iterator is expressed as follows:
[0018]
[0019]
[0020]
[0021] Where is the numerical implementation of the Lagrangian function at the t-th iteration, is the n-th realization of the achievable rate of the k-th user at the t-th iteration, N b is the batch data size, K is the number of the users, is the effective bandwidth of the k-th user, θ k is a constant that can be obtained, E g is to take the expectation of the small-scale channel gain, ω(t) is the parameter of the power neural network at the t-th iteration, ω(t + 1) is the parameter of the power neural network at the (t + 1)-th iteration, and φ(t) is the learning rate at the t-th iteration, is the vector differential operator, W k is the broadband of the k-th user, W k (t) is the broadband of the k-th user at the t-th iteration, W k(t + 1) is the bandwidth of the k-th user at the (t + 1)-th iteration, λ k is the long-term constraint multiplier of the k-th user, λ k (t) is the long-term constraint multiplier of the k-th user at the t-th iteration, λ k (t + 1) is the long-term constraint multiplier of the k-th user at the (t + 1)-th iteration.
[0022] In some embodiments of the present application, the resource allocation according to the parameters of the current moment solver and the small-scale channel gain at the current moment includes:
[0023] According to the power neural network parameters among the parameters of the current moment solver and the small-scale channel gain at the current moment, obtain the power of each user at the current moment;
[0024] Determine the user bandwidth among the parameters of the current moment solver as the bandwidth of each user at the current moment;
[0025] Perform resource allocation according to the bandwidth and power of each user at the current moment.
[0026] In some embodiments of the present application, the model used when the iterator first performs online training is an iterator model pre-trained offline; wherein, the parameters in the offline pre-trained iterator model are used as the initial parameters of the solver when the iterator is first online trained.
[0027] According to the second aspect of the present application, there is provided a machine learning model training and inference method adapted to dynamic changes in the communication environment, including:
[0028] Obtain slow-changing parameters and fast-changing parameters in the offline environment;
[0029] Pre-train a preset iterator model using unsupervised deep learning techniques, the slow-changing parameters and the fast-changing parameters in the offline environment;
[0030] Obtain the slow-changing parameters at the current moment in the dynamic environment and the fast-changing parameters at the current moment;
[0031] Online train the iterator according to the slow-changing parameters at the current moment in the dynamic environment, the parameters of the solver at the previous moment, and the fast-changing parameters in the offline environment to update the parameters of the solver at the current moment; wherein, the model used when the iterator first performs online training is an iterator model pre-trained offline, and the parameters in the offline pre-trained iterator model are used as the initial parameters of the solver when the iterator is first online trained;
[0032] Perform real-time inference according to the parameters of the current moment solver and the fast-changing parameters at the current moment.
[0033] According to the third aspect of the present application, a network-side device is provided, including a memory, a transceiver, and a processor; wherein,
[0034] The memory is used to store computer programs; the transceiver is used to send and receive data under the control of the processor; the processor is used to read the computer programs in the memory and perform the following operations:
[0035] Obtain the large-scale channel gain at the current moment and the small-scale channel gain at the current moment in the dynamic environment;
[0036] Obtain the target value of the small-scale channel gain, and obtain the parameters of the solver at the previous moment; the parameters include user bandwidth, multiplier, and power neural network parameters;
[0037] Online train the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment;
[0038] Perform resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment.
[0039] In some embodiments of the present application, the obtaining the target value of the small-scale channel gain includes:
[0040] Obtain the small-scale channel gain from the offline environment, and use the small-scale channel gain obtained from the offline environment as the target value of the small-scale channel gain.
[0041] In some embodiments of the present application, the online training the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment includes:
[0042] Calculate the corresponding gradient according to the large-scale channel gain at the current moment;
[0043] Use the parameters of the solver at the previous moment as the initial value of the iterator at the current moment;
[0044] Perform N iterations on the iterator at the current moment according to the corresponding gradient and the target value of the small-scale channel gain;
[0045] Use the parameters obtained after the N iterations of the iterator at the current moment as the parameters of the solver at the current moment.
[0046] In the embodiments of the present application, the formula of the iterator is expressed as follows:
[0047]
[0048]
[0049]
[0050] Where is the numerical implementation of the Lagrangian function at the t-th iteration, is the n-th implementation of the achievable rate of the k-th user at the t-th iteration, N b is the batch data size, K is the number of the users, is the effective bandwidth of the k-th user, θ k is a constant that can be obtained, E g is to take the expectation of the small-scale channel gain, ω(t) is the parameter of the power neural network at the t-th iteration, ω(t + 1) is the parameter of the power neural network at the (t + 1)-th iteration, and φ(t) is the learning rate at the t-th iteration, is the vector differential operator, W k is the broadband of the k-th user, W k (t) is the broadband of the k-th user at the t-th iteration, W k (t + 1) is the broadband of the k-th user at the (t + 1)-th iteration, λ k is the long-term constraint multiplier of the k-th user, λ k (t) is the long-term constraint multiplier of the k-th user at the t-th iteration, λ k (t + 1) is the long-term constraint multiplier of the k-th user at the (t + 1)-th iteration.
[0051] In some embodiments of the present application, the resource allocation based on the parameters of the solver at the current moment and the small-scale channel gain at the current moment includes:
[0052] Obtain the power of each user at the current moment according to the power neural network parameters among the parameters of the solver at the current moment and the small-scale channel gain at the current moment;
[0053] Determine the user broadband among the parameters of the solver at the current moment as the broadband of each user at the current moment;
[0054] Perform resource allocation according to the bandwidth and power of each user at the current moment.
[0055] In some embodiments of the present application, the model used by the iterator during the first online training is an iterator model pre-trained offline; wherein, the parameters in the offline pre-trained iterator model are used as the initial parameters of the solver when the iterator is first online trained.
[0056] According to a fourth aspect of the present application, there is provided a network-side device, including a memory, a transceiver, and a processor; wherein,
[0057] The memory is used to store a computer program; the transceiver is used to transmit and receive data under the control of the processor; the processor is used to read the computer program in the memory and perform the following operations:
[0058] Obtain slow-varying parameters and fast-varying parameters in the offline environment;
[0059] Pre-train a preset iterator model using unsupervised deep learning technology, the slow-varying parameters and fast-varying parameters in the offline environment;
[0060] Obtain the slow-varying parameters at the current moment and the fast-varying parameters at the current moment in the dynamic environment;
[0061] Online train the iterator according to the slow-varying parameters at the current moment in the dynamic environment, the parameters of the solver at the previous moment, and the fast-varying parameters in the offline environment to update the parameters of the solver at the current moment; wherein, the model used by the iterator during the first online training is an iterator model pre-trained offline, and the parameters in the offline pre-trained iterator model are used as the initial parameters of the solver when the iterator is first online trained;
[0062] Perform real-time inference according to the parameters of the solver at the current moment and the fast-varying parameters at the current moment.
[0063] According to a fifth aspect of the present application, there is provided a wireless resource allocation device, including:
[0064] A first acquisition unit, configured to acquire the large-scale channel gain at the current moment and the small-scale channel gain at the current moment in the dynamic environment;
[0065] A second acquisition unit, configured to acquire the target value of the small-scale channel gain;
[0066] A third acquisition unit, configured to acquire the parameters of the solver at the previous moment; the parameters include user bandwidth, multiplier, and power neural network parameters;
[0067] An online training unit, configured to online train the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment;
[0068] A resource allocation unit, configured to perform resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment.
[0069] In some embodiments of the present application, the second acquisition unit is specifically configured to: acquire the small-scale channel gain from an offline environment, and use the small-scale channel gain acquired from the offline environment as the target value of the small-scale channel gain.
[0070] In some embodiments of the present application, the online training unit is specifically configured to: calculate the corresponding gradient according to the large-scale channel gain at the current moment; use the parameters of the solver at the previous moment as the initial value of the iterator at the current moment; perform N iterations on the iterator at the current moment according to the corresponding gradient and the target value of the small-scale channel gain; and use the parameters obtained after the N iterations of the iterator at the current moment as the parameters of the solver at the current moment.
[0071] In an embodiment of the present application, the formula of the iterator is expressed as follows:
[0072]
[0073]
[0074]
[0075] Where is the numerical implementation of the Lagrangian function at the t-th iteration, is the n-th realization of the achievable rate of the k-th user at the t-th iteration, N b is the batch data size, K is the number of the users, is the effective bandwidth of the k-th user, θ k is a constant that can be obtained, E g is to take the expectation of the small-scale channel gain, ω(t) is the parameter of the power neural network at the t-th iteration, ω(t + 1) is the parameter of the power neural network at the (t + 1)-th iteration, and φ(t) is the learning rate at the t-th iteration, is the vector differential operator, W k is the broadband of the k-th user, W k (t) is the broadband of the k-th user at the t-th iteration, W k (t + 1) is the broadband of the k-th user at the (t + 1)-th iteration, λ k is the long-term constraint multiplier of the k-th user, λ k (t) is the long-term constraint multiplier of the k-th user at the t-th iteration, λ k(t + 1) is the long-term constraint multiplier of the k-th user at the (t + 1)-th iteration.
[0076] In some embodiments of the present application, the resource allocation unit is specifically configured to: obtain the power of each user at the current moment according to the power neural network parameters among the parameters of the current moment solver and the small-scale channel gain at the current moment; determine the user bandwidth among the parameters of the current moment solver as the bandwidth of each user at the current moment; and perform resource allocation according to the bandwidth and power of each user at the current moment.
[0077] In some embodiments of the present application, the model used by the iterator for the first online training is an iterator model pre-trained offline; wherein, the parameters in the offline pre-trained iterator model are used as the initial parameters of the solver when the iterator is first online trained.
[0078] According to a sixth aspect of the present application, there is provided a machine learning model training and inference device adapted to dynamic changes in the communication environment, including:
[0079] A first acquisition unit for acquiring slow-changing parameters and fast-changing parameters in an offline environment;
[0080] A pre-training unit for pre-training a preset iterator model using unsupervised deep learning techniques, the slow-changing parameters and fast-changing parameters in the offline environment;
[0081] A second acquisition unit for acquiring the slow-changing parameters at the current moment and the fast-changing parameters at the current moment in a dynamic environment;
[0082] An online training unit for online training an iterator according to the slow-changing parameters at the current moment in the dynamic environment, the parameters of the solver at the previous moment, and the fast-changing parameters in the offline environment to update the parameters of the solver at the current moment; wherein, the model used by the iterator for the first online training is an iterator model pre-trained offline, and the parameters in the offline pre-trained iterator model are used as the initial parameters of the solver when the iterator is first online trained;
[0083] An inference unit for performing real-time inference according to the parameters of the solver at the current moment and the fast-changing parameters at the current moment.
[0084] According to a seventh aspect of the present application, there is provided a processor-readable storage medium storing a computer program for causing the processor to execute the wireless resource allocation method described in the first aspect of the present application, or execute the machine learning model training and inference method for adapting to dynamic changes in the communication environment described in the second aspect of the present application.
[0085] According to the technical solution of the embodiment of the present application, by obtaining the large-scale channel gain at the current moment in the dynamic environment and the small-scale channel gain at the current moment, obtaining the target value of the small-scale channel gain, and obtaining the parameters of the solver at the previous moment, training the online training iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment, and performing resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment, the OFDMA network resources can be reasonably allocated to different users, thereby enhancing the performance of the entire network. In addition, when the present application performs online training on the iterator, only the small-scale channel gain in the environmental parameters needs to be learned and the large-scale channel gain needs to be tracked, so that the neural network structure in online training is simpler than that in offline training, the real-time decision-making speed is faster, and the adaptability to the environment is stronger. Description of the Drawings
[0086] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0087] Figure 1 It is a flowchart of a machine learning model training and inference method for adapting to dynamic changes in the communication environment provided by an embodiment of the present application.
[0088] Figure 2 It is an implementation flowchart of the online training method provided by an embodiment of the present application.
[0089] Figure 3 It is an implementation flowchart of online training in the O-RAN architecture according to an embodiment of the present application.
[0090] Figure 4 It is a schematic diagram of the time scale of channel change and frame structure provided by an embodiment of the present application.
[0091] Figure 5 It is an implementation flowchart of real-time resource allocation for URLLC provided by an embodiment of the present application.
[0092] Figure 6 It is a flowchart of a wireless resource allocation method provided by an embodiment of the present application.
[0093] Figure 7 It is a structural block diagram of a network-side device provided by an embodiment of the present application.
[0094] Figure 8 It is a structural block diagram of a wireless resource allocation device provided by an embodiment of the present application.
[0095] Figure 9 It is a structural block diagram of a machine learning model training and inference device provided by an embodiment of the present application, which adapts to dynamic changes in the communication environment. Detailed implementation manners
[0096] In the embodiments of the present invention, the term "and / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0097] In the embodiments of the present application, the term "a plurality of" refers to two or more, and other quantifiers are similar thereto.
[0098] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0099] The network-side device involved in the embodiments of this application can be a base station, which can include multiple cells that provide services to terminals. Depending on the specific application scenario, the base station can also be referred to as an access point, or it can be a device in the access network that communicates with wireless terminal devices through one or multiple sectors over the air interface, or have other names. The network-side device can be used to mutually replace the received air frames and Internet Protocol (IP) packets, acting as a router between the wireless terminal device and the rest of the access network, where the rest of the access network can include an Internet Protocol (IP) communication network. The network-side device can also coordinate the attribute management of the air interface. For example, the network-side device involved in the embodiments of this application can be a network-side device (Base Transceiver Station, BTS) in a Global System for Mobile communications (GSM) or Code Division Multiple Access (CDMA), or a network-side device (NodeB) in a Wide-band Code Division Multiple Access (WCDMA), or an evolved network-side device (evolutional Node B, eNB or e-NodeB) in a Long Term Evolution (LTE) system, a 5G base station (gNB) in a 5G network architecture (next generation system), or a Home evolved Node B (HeNB), a relay node, a femto, a pico, etc. This is not limited in the embodiments of this application. In some network architectures, the network-side device can include a centralized unit (centralized unit, CU) node and a distributed unit (distributed unit, DU) node, and the centralized unit and the distributed unit can also be arranged separately geographically.
[0100] It should be noted that wireless resource allocation belongs to a complex optimization problem in wireless communication. Facing the complex optimization problems in wireless communication, traditional optimization methods are difficult to obtain accurate optimization results in a short time (such as 1 millisecond). To overcome this difficulty, some researchers have proposed a method of using neural networks to learn the optimization results to quickly obtain the optimization results. This method aims at parameter optimization problems, solves the optimization problems offline through traditional optimization methods in various environmental states, and uses supervised learning to train the neural network to fit the relationship between the environmental state and the optimization results. Finally, the fully trained neural network is used for online decision-making to give the optimization results in real time according to the actual environment. For problems that are difficult to solve using traditional optimization methods, such as functional optimization problems, unsupervised learning can be used for training. However, limited by the scale of xunlji and the training time, the generalization ability of the neural network to environmental changes is limited. When the actual environment is significantly different from the training environment, the neural network cannot provide reliable results.
[0101] Based on the above problems, this application designs a machine learning architecture based on transfer learning by taking the unsupervised deep learning method as an example, that is, proposes a machine learning model training and inference method that can adapt to the dynamic changes of the communication environment, combines offline learning and online learning, and can be applied to the machine learning of O-RAN and other systems. As Figure 1 shown, the machine learning model training and inference method that adapts to the dynamic changes of the communication environment in the embodiments of this application is applied to the network-side device, and may include the following steps.
[0102] In step S101, slow-changing parameters and fast-changing parameters in the offline environment are obtained.
[0103] In step S102, the preset iterator model is pre-trained using unsupervised deep learning technology, slow-changing parameters, and fast-changing parameters in the offline environment.
[0104] In step S103, the slow-changing parameters at the current moment and the fast-changing parameters at the current moment in the dynamic environment are obtained.
[0105] In step S104, the iterator is online trained according to the slow-changing parameters at the current moment in the dynamic environment, the parameters of the solver at the previous moment, and the fast-changing parameters in the offline environment to update the parameters of the solver at the current moment.
[0106] Among them, in the embodiments of this application, the model used when the iterator is first online trained is the iterator model pre-trained offline, and the parameters in the iterator model pre-trained offline are used as the initial parameters of the solver when the iterator is first online trained.
[0107] In step S105, real-time inference is performed according to the parameters of the solver at the current moment and the fast-changing parameters at the current moment.
[0108] The following will combine Figure 2 and Figure 3 to describe in detail the implementation process of the machine learning model training and inference method that adapts to the dynamic changes of the communication environment in the embodiments of the present application. First, it should be noted that as Figure 2 shown, the machine learning model training and inference method that adapts to the dynamic changes of the communication environment in the embodiments of the present application mainly includes the following three steps:
[0109] 1. Offline pre-training: Perform pre-training through unsupervised learning, and the training data comes from an offline environment. Here, the offline environment refers to a data set obtained and stored in advance under non-real-time conditions, so as to obtain an initialized solver (wherein, the solver can be understood as an optimization parameter) and an iterator model, and download the solver and the iterator model to the online training platform.
[0110] 2. Online training: Track the slow-changing parameters in the dynamic environment. Based on the idea of transfer learning, use the parameters in the solver at the previous moment as the initial value of the iterator at the current moment. Input a slow-changing parameter taken from the dynamic environment into the iterator, and at the same time use the fast-changing parameters in the offline environment for iteration to update the parameters in the solver.
[0111] 3. ML inference model: Use the iterator model after online training, utilize the results of the updated solver, and perform real-time inference (such as real-time resource allocation) based on the current slow-changing parameters and fast-changing parameters.
[0112] In some embodiments of the present application, when this method is applied to the O-RAN (Open RAN) architecture, this part of the pre-training can be implemented in a Non Real Time RAN Intelligent Controller (Non-RT RIC), and these two parts of online tracking and ML model inference can be implemented in a NearRealTime RAN Intelligent Controller (Near-RT RIC).
[0113] It should be noted that before training the machine learning model, it is necessary to first define the problem model. For example, consider a general resource optimization problem that minimizes the long-term loss and maximizes the profit while satisfying long-term constraints and short-term constraints, in the following form:
[0114]
[0115]
[0116] g(x, f(θ q ),θq , θ s ) ≤ 0 #(1b)
[0117] This problem is a functional optimization problem, where f(θ q ) is a control function dependent on the fast-varying parameter θ q , x is a decision vector independent of θ q , and the optimization results of both are affected by the slow-varying parameter θ s . J, c, and g are the evaluation indicators corresponding to the long-term objective, long-term constraint, and short-term constraint, respectively, and all depend on the policy x and f(θ q ) as well as the parameters θ q and θ s . For simplicity, the independent variables of the three evaluation indicators are omitted in the following representations.
[0118] In the offline pre-training stage, for any given slow-varying parameter θ s , to solve problem (1), this application adopts an unsupervised deep learning method. First, by introducing Lagrange multipliers, the original constrained problem (1) can be transformed into the following form of the primal-dual problem:
[0119]
[0120] s.t. λ ≥ 0, μ(θ q ) ≥ 0 #(2a)
[0121] where λ is the multiplier vector of the long-term constraint, and μ(θ q ) is the multiplier function of the instantaneous constraint. To transform the undetermined function into the parameter space for solution, a neural network is used to parameterize the control function f(θ q ) and the multiplier function μ(θ q ):
[0122]
[0123] where functions such as ReLU and Softplus can be selected as the activation function of the output layer to ensure that its output result is non-negative.
[0124] Subsequently, the Lagrangian function L is used as the loss function, and stochastic gradient descent is used to optimize the variable x and train the policy network parameter ω f , and stochastic gradient ascent is used to optimize the multiplier λ and train the multiplier network parameter ω μ :
[0125]
[0126]
[0127]
[0128]
[0129] where + = max{z, 0}, is the numerical implementation of the Lagrangian function at the t-th iteration. The trained network can then give the decision result in real time according to the dynamically changing θ q However, when θ s changes, retraining is required.
[0130] For the O-RAN system, as Figure 3 shown, offline training is performed on the Non-RT RIC, and the training data comes from the O1 interface, including fast-changing parameters and slow-changing parameters. The offline-trained model is downloaded to the Near-RT RIC through the O1 interface as the iterative model, and at the same time, the optimized parameters obtained from the offline training are sent to the Near-RT RIC through the O1 interface as the initial parameters of the solver.
[0131] In the first step of the implementation process, pre-training is performed through the offline training method, and the solver parameters include the optimization variable x, the multiplier λ, and the neural network parameters ω f and ω μ .
[0132] In the second step of the implementation process, the dynamic environment parameters are divided into fast-changing parameters and slow-changing parameters. Taking the channel gain as an example, the average of the channel samples within a period T L is taken as the slow-changing parameter for this period, and all the samples within this period minus the average value are taken as the fast-changing parameters, that is:
[0133]
[0134] θ q i = θ i - θ s
[0135] where N represents the total number of samples within T L , θ i represents the i-th channel sample, and θ q i represents the i-th fast-changing parameter. The slow-changing parameter will be used in the iterative process of online training, and the fast-changing parameter here will be used for real-time decision-making.
[0136] In this application, iteration is performed according to the above formulas (3a - 3d), and each iteration is based on the latest observation of the slow-changing parameter θ s in the dynamic environment. For example, for the l-th observation, θs The value of is denoted as Substitute into q and calculate the corresponding gradient, and then perform N iterations using the fast-changing parameter. Since in the above problem model, the mapping relationship of f(θ q ) is determined by the distribution of θ q , rather than a specific value of θ
[0137] For the (l + 1)-th observation, θ s changes from to Based on the idea of transfer learning in this application, since θ s changes slowly and the difference between adjacent observations is small, the optimal strategy in the new environment and the optimal strategy in the original environment often differ little. Therefore, the parameters obtained from the previous optimization and training can be used as the initial value for policy optimization in the new environment, that is
[0138] x l+1 (1) = x l (N)
[0139]
[0140] λ l+1 (1) = λ l (N)
[0141]
[0142] Thus, in each round of training, only a small number of iterations are required to correct the previous policy, obtain the optimal network parameters and the optimal solution at the current moment, so as to achieve the purpose of online tracking.
[0143] For the O-RAN system, the solver and iterator models obtained from offline training are downloaded to the Near-RTRIC through the O1 interface for online training. The input data of the iterator model for online training includes the fast-changing parameters of the offline environment obtained from the O1 interface, the slow-changing parameters of the dynamic environment obtained from the E2 interface, and the parameters in the solver at the previous moment (used as the initial value of the iterator at the current moment). Moreover, the parameters of the solver obtained by each iterator during online training can be used as the initial value of the solver for the next online training.
[0144] The iteration of the online iterator is determined by equations (3a - 3d), and the initial iteration value comes from equation (4); while the iteration of the iterator in the offline pre-training method is determined by equations (3a - 3d), and the initial iteration value is random.
[0145] In the third step of the implementation process, this application is based on the current environmental parameter θ s obtained below to perform real-time inference on f(θ q ) and update the optimization variable x in real time based on the iterative relation (3a). For the O-RAN system, the model for real-time inference comes from the iterative relation (3a), and the input parameters of this model are the slow-changing parameters and fast-changing parameters in the dynamic environment from the E2. Among them, for the first online training, its parameters need to be initialized through the pre-training in the first step.
[0146] In this application, the fast-changing parameters are generalized as the input of the neural network, and the slow-changing parameters can be dynamically tracked through online training. This can simplify the neural network structure of online unsupervised learning on the one hand and can well adapt to environmental changes on the other hand.
[0147] It can be seen that based on unsupervised deep learning and transfer learning, this application provides a method for training and inferring a machine learning model that can adapt to the dynamic changes of the communication environment. It makes full use of the characteristics of the fast and slow parameters in the offline environment and the dynamic environment, as well as the iterator, to reconstruct the online training system of the O-RAN machine learning model. Utilizing the temporal correlation between the previous and current moments in the real-time scenario, taking the optimal solution of the neural network parameters in the previous moment as the initial value of the neural network in the next moment greatly reduces the number of iterations for each training. In this way, it can not only track the dynamic changes of the slow-changing parameters online but also make real-time decisions according to the current fast-changing parameters, thus ensuring the generalization performance and real-time performance of the model. And unless the machine learning model of the offline training is changed, there is no need to download the machine learning model to the online training system after the offline training is completed, but only the relevant parameters of the offline training need to be downloaded, thereby improving the efficiency of the entire training system.
[0148] It can be understood that for a traditional O-RAN ML model used in real-time scenarios, one method is to directly use the offline model in the online mode for online training and inference. When the environment changes, the model is retrained from scratch, which brings huge training complexity and is difficult to meet the real-time requirements. Another method is to learn as many offline environmental parameters as possible, but this will increase the scale and structure of the neural network, resulting in problems such as difficult training and long decision-making time. Compared with the above traditional offline training methods, the online training of this application only needs to learn the fast-changing parameters in the environmental parameters and track the slow-changing parameters. This makes the neural network structure in the online training method simpler than that in offline training, with a faster real-time decision-making speed and stronger adaptability to the environment.
[0149] This application provides an example of wireless resource allocation. Consider an orthogonal frequency division multiple access (OFDMA) downlink for ultra-reliable low-latency communication (URLLC) services. There is a base station with a maximum transmit power of P max in a cell, equipped with N t antennas, serving K single-antenna users.
[0150] The time scale division in this example is as Figure 4 shown: The correlation time of the small-scale channel gain is T s , and each T s is further divided into multiple frame structures. The length of each frame is T f . To meet the stringent latency requirements of URLLC, frequency hopping technology is considered, so that the small-scale channel gains between each frame are independent. The uplink transmission and downlink transmission are completed within one frame, where the downlink transmission time is τ. Additionally, it is assumed that the large-scale channel gain is approximately constant within T L .
[0151] In URLLC, due to the short transmission time, the code length of channel coding is short, and the impact of decoding errors on reliability cannot be ignored. Therefore, this application considers the achievable rate under the finite blocklength regime. In a quasi-static flat fading channel, when the channel state information is known at both the transmitter and the receiver, the achievable rate (in packets / frame) of the k-th user can be estimated by the following formula:
[0152]
[0153] where is the decoding error probability of the k-th user, α k and gk is the large-scale channel gain and small-scale channel gain of the k-th user, and N0 is the one-sided power spectral density of the noise. is the reciprocal of the Gaussian Q function, V k is a parameter representing channel dispersion. When the signal-to-noise ratio is large, V k ≈ 1.
[0154] To utilize multi-user diversity, this application allocates the transmission power of each user according to the small-scale channel gains g = {g1, g2, …, g K}. At the same time, to reduce the computational complexity, it is assumed that the bandwidth allocation result only depends on the large-scale channel gain α. Combining the foregoing time-scale description, here g is a fast-varying parameter and α is a slow-varying parameter. The optimization objective of this application is to minimize the total user bandwidth under the premise of meeting the QoS requirements of URLLC. The optimization variables are the bandwidth and power allocated by the base station to each user:
[0155]
[0156]
[0157]
[0158] In the above formula, (5a) represents the constraint relationship characterizing the QoS requirements of URLLC by introducing the effective capacity and the effective bandwidth . The left side of the inequality is the expression of the effective capacity , and θ k is a constant that can be obtained, is to take the expectation of the small-scale channel gain. In addition, when considering that the packet arrival process is a Poisson process, where a k is the average packet arrival rate of k users. (5a) is a long-term constraint, indicating the limitations that the delay and reliability of users need to meet within the time T L . And (5b) is a short-term constraint, indicating that the maximum power limit and non-negativity limit need to be met within each T f .
[0159] Under the constraints of the above two time scales, problem (5) becomes a functional optimization problem, which is difficult to solve with traditional optimization algorithms (such as convex optimization). This application will solve this problem based on the offline learning optimization method and the online training method respectively.
[0160] In the offline training stage, first introduce the Lagrange multiplier to transform the original problem with constraints (5) into the following form of the dual problem:
[0161]
[0162] such that P k (g)≥0, W k ≥0, h(g)≥0, λ k ≥0 (6)
[0163] where λ k is the long - term constraint multiplier of the k - th user, and h(g) is the multiplier function of the instantaneous constraint. The power function P(g) and the multiplier function h(g) are estimated parametrically using a neural network:
[0164]
[0165] In this example, the output layer of the neural network is chosen as Softmax, so that the maximum power limit in (5b) can actually be automatically satisfied, and at this time, it is not necessary to parameterize h(g) anymore.
[0166] For each fixed slow - varying parameter α, the optimal resource allocation can be iteratively solved based on the following formula:
[0167]
[0168]
[0169]
[0170] where is the n - th realization of the achievable rate of the k - th user at the t - th iteration, and N b is the batch data size. Note that when α changes, it is necessary to re - iteratively solve problem (6) from the beginning using formula (7).
[0171] In the online training phase, the process of the base station performing real - time resource allocation is as Figure 5 shown:
[0172] 1. The base station initializes the solver through pre - training by an offline training method. The solver parameters include the user bandwidth multiplier and the power neural network where N0 represents the number of iterations of unsupervised learning, and N0 >> N.
[0173] 2. Track the large - scale channel gain in the dynamic environment. Among them, formula (7) is also used for solving, but the difference is that each solution is based on the latest observation of α at the current time. For the observation at the l - th moment, the value of α is denoted as α l , substitute α = α l into Calculate the corresponding gradients and then perform N iterations according to Equation (7). For the observation at the (l + 1)-th moment, α changes from α l to α l+1 . Since α changes slowly and the difference between adjacent observations is small, the optimal strategy in the new environment often differs little from the optimal strategy in the original environment. Therefore, the parameters obtained from the previous optimization and training can be used as the initial values for strategy optimization in the new environment, that is:
[0174] W l+1 (1) = W l (N)
[0175]
[0176] λ l+1 (1) = λ l (N)#(8)
[0177] Thus, in each round of training, only a small number of iterations are required to correct the previous strategy to obtain the optimal network parameters and the optimal solution at the current moment, thereby achieving the purpose of online tracking. Input a large-scale channel gain into the iterator, and at the same time use N small-scale channel gains / batches to iterate N times to update the parameters in the solver at the current moment l + 1. Note that the small-scale channel gains used for iteration do not need to be real-time sampled and can be data from an offline environment.
[0178] 3. The base station uses the solver at the current moment l + 1 to perform real-time resource allocation based on the small-scale channel gain g at the current moment l+1 .
[0179] For example, as Figure 6 shown, the wireless resource allocation method of the embodiment of the present application is applied to a network-side device and may include the following steps.
[0180] In step S601, obtain the large-scale channel gain at the current moment in the dynamic environment and the small-scale channel gain at the current moment.
[0181] In step S602, obtain the target value of the small-scale channel gain and obtain the parameters of the solver at the previous moment; the parameters include user bandwidth, multiplier, and power neural network parameters.
[0182] In the embodiments of the present application, the small-scale channel gain can be obtained from the offline environment, and the small-scale channel gain obtained from the offline environment is used as the target value of the small-scale channel gain. That is to say, the small-scale channel gain used in the online training of the iterator does not need to be real-time sampled and can be data from the offline environment. Thus, by adopting the small-scale channel gain in the offline environment, sufficient samples can be directly generated according to the distribution, and there is no need to wait for too long. Therefore, the adoption of offline samples can balance the learning performance and real-time requirements.
[0183] In step S603, the iterator is online trained according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment.
[0184] In some embodiments of the present application, the corresponding gradient can be calculated according to the large-scale channel gain at the current moment; the parameters of the solver at the previous moment are used as the initial value of the iterator at the current moment; according to the corresponding gradient and the target value of the small-scale channel gain, the iterator at the current moment is iterated N times; the parameters obtained after the iterator at the current moment is iterated N times are used as the parameters of the solver at the current moment. Among them, the formula representation of the iterator can be as follows:
[0185]
[0186]
[0187]
[0188] Among them is the numerical implementation of the Lagrangian function at the t-th iteration, is the n-th realization of the achievable rate of the k-th user at the t-th iteration, N b is the batch data size, K is the number of users, is the effective bandwidth of the k-th user, θ k is a constant that can be obtained, E g is to take the expectation of the small-scale channel gain, ω(t) is the parameter of the power neural network at the t-th iteration, ω(t + 1) is the parameter of the power neural network at the t + 1-th iteration, and φ(t) is the learning rate at the t-th iteration, is the vector differential operator, W k is the broadband of the k-th user, W k (t) is the broadband of the k-th user at the t-th iteration, W k (t + 1) is the broadband of the k-th user at the t + 1-th iteration, λ k is the long-term constraint multiplier of the k-th user, λ k$(t)$ is the long-term constraint multiplier of the $k$-th user at the $t$-th iteration, $\lambda$ k $(t + 1)$ is the long-term constraint multiplier of the $k$-th user at the $(t + 1)$-th iteration.
[0189] It should be noted that the model used by the iterator for the first online training is the iterator model pre-trained offline; among them, the parameters in the iterator model pre-trained offline are used as the initial parameters of the solver when the iterator is first trained online.
[0190] In step S604, resource allocation is performed according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment.
[0191] For example, as Figure 5 shown, when the iterator is trained online according to the large-scale channel gain $\alpha$ l+1 at the current moment, the parameters of the solver at the previous moment (i.e., $\omega$ l , $W$ k l , $\lambda$ k l ) and the small-scale channel gain target value (i.e., the small-scale channel gain $g$ in the offline environment) to update the parameters of the solver at the current moment $l + 1$ (i.e., $\omega$ l+1 , $W$ k l+1 , $\lambda$ k l+1 ), the power of each user at the current moment can be calculated using the power calculation formula l+1 according to the power neural network parameter $\omega$ among the parameters of the solver at the current moment $l + 1$ l+1 and the small-scale channel gain $g$ at the current moment $l + 1$. The user bandwidth $W$ among the parameters of the solver at the current moment $l + 1$ is determined as the bandwidth of each user at the current moment $l + 1$. Resource allocation is performed according to the bandwidth and power of each user at the current moment $l + 1$. k l+1
[0192] In summary, the wireless resource allocation method according to the embodiments of the present application obtains the large-scale channel gain and the small-scale channel gain at the current moment in a dynamic environment, obtains the target value of the small-scale channel gain, and obtains the parameters of the solver at the previous moment. The iterator is online trained according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment. Moreover, resource allocation is performed according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment, so that the OFDMA network resources can be reasonably allocated to different users, thereby enhancing the performance of the entire network. In addition, when the present application performs online training on the iterator, only the small-scale channel gain in the environmental parameters needs to be learned, and the large-scale channel gain is tracked, making the neural network structure in online training simpler than that in offline training, with a faster real-time decision-making speed and a stronger adaptability to the environment.
[0193] Figure 7 is a structural block diagram of a network-side device according to an embodiment of the present application. As Figure 7 shown, the network-side device may include: a memory 701, a transceiver 702, and a processor 703, and data transmission is realized through a bus interface.
[0194] Among them, the memory 701 is used to store computer programs;
[0195] The transceiver 702 is used to transmit and receive data under the control of the processor 703; the transceiver 702 may be multiple components, that is, including a transmitter and a receiver, and provides a unit for communicating with various other devices on a transmission medium, and these transmission media include wireless channels, wired channels, optical fiber cables, and other transmission media;
[0196] The processor 703 is used to read the computer program in the memory and execute corresponding operations. The processor 703 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD), and the processor may also adopt a multi-core architecture.
[0197] In the embodiments of the present application, the processor 703 reads the computer program stored therein and performs the following operations:
[0198] S701', obtain the large-scale channel gain and the small-scale channel gain at the current moment in a dynamic environment;
[0199] S702’, obtain the target value of the small-scale channel gain and obtain the parameters of the solver at the previous moment; the parameters include the user bandwidth, multiplier, and power neural network parameters;
[0200] In the embodiments of the present application, the small-scale channel gain can be obtained from the offline environment, and the small-scale channel gain obtained from the offline environment is used as the target value of the small-scale channel gain.
[0201] S703’, online train the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment;
[0202] In some embodiments, calculate the corresponding gradient according to the large-scale channel gain at the current moment; use the parameters of the solver at the previous moment as the initial value of the iterator at the current moment; perform N iterations on the iterator at the current moment according to the corresponding gradient and the target value of the small-scale channel gain; use the parameters obtained after N iterations of the iterator at the current moment as the parameters of the solver at the current moment.
[0203] In the embodiments of the present application, the formula of the iterator is expressed as follows:
[0204]
[0205]
[0206]
[0207] where is the numerical implementation of the Lagrangian function at the t-th iteration, is the n-th realization of the achievable rate of the k-th user at the t-th iteration, N b is the batch data size, K is the number of users, is the effective bandwidth of the k-th user, θ k is a constant that can be obtained, E g is to take the expectation of the small-scale channel gain, ω(t) is the parameter of the power neural network at the t-th iteration, ω(t + 1) is the parameter of the power neural network at the (t + 1)-th iteration, φ(t) is the learning rate at the t-th iteration, is the vector differential operator, W k is the bandwidth of the k-th user, W k (t) is the bandwidth of the k-th user at the t-th iteration, W k (t + 1) is the bandwidth of the k-th user at the (t + 1)-th iteration, λ k is the long-term constraint multiplier of the k-th user, λ k$(t)$ is the long-term constraint multiplier of the $k$-th user at the $t$-th iteration, $\lambda$ k $(t + 1)$ is the long-term constraint multiplier of the $k$-th user at the $(t + 1)$-th iteration.
[0208] It should be noted that, in the embodiments of the present application, the model used by the iterator for the first online training is an iterator model pre-trained offline; among them, the parameters in the iterator model pre-trained offline are used as the initial parameters of the solver when the iterator is first online trained.
[0209] S704’, perform resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment.
[0210] In some embodiments, the power of each user at the current moment can be obtained according to the power neural network parameters among the parameters of the solver at the current moment and the small-scale channel gain at the current moment; the user bandwidth among the parameters of the solver at the current moment is determined as the bandwidth of each user at the current moment; resource allocation is performed according to the bandwidth and power of each user at the current moment.
[0211] The network-side device in the embodiments of the present application can reasonably allocate OFDMA network resources to different users by obtaining the large-scale channel gain and the small-scale channel gain at the current moment in the dynamic environment, obtaining the target value of the small-scale channel gain, obtaining the parameters of the solver at the previous moment, online training the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment, and performing resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment, thereby enhancing the performance of the entire network. In addition, when the present application performs online training on the iterator, it only needs to learn the small-scale channel gain in the environmental parameters and track the large-scale channel gain, making the neural network structure in the online training simpler than that in the offline training, with a faster real-time decision-making speed and a stronger adaptability to the environment.
[0212] In other embodiments of the present application, the processor 703 reads the computer program stored therein and performs the following operations:
[0213] S71’, obtain the slow-varying parameters and fast-varying parameters in the offline environment;
[0214] S72’, pre-train a preset iterator model using unsupervised deep learning technology, the slow-varying parameters and fast-varying parameters in the offline environment;
[0215] S73’, obtain the slow-varying parameters and the fast-varying parameters at the current moment in the dynamic environment;
[0216] S74’, online train the iterator based on the slow-changing parameters at the current moment in the dynamic environment, the parameters of the solver at the previous moment, and the fast-changing parameters in the offline environment, so as to update the parameters of the solver at the current moment; wherein, the model used when the iterator is first online trained is the iterator model pre-trained offline, and the parameters in the pre-trained offline iterator model are used as the initial parameters of the solver when the iterator is first online trained;
[0217] S75’, perform real-time inference according to the parameters of the solver at the current moment and the fast-changing parameters at the current moment.
[0218] It should be noted here that the above network-side device provided by the embodiments of the present invention can implement all the method steps implemented by the above method embodiments, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described in this embodiment.
[0219] According to the network-side device of the embodiments of the present application, by utilizing the time correlation between the previous and current moments in the real-time scenario, the optimal solution of the neural network parameters in the previous moment is used as the initial value of the neural network in the next moment, which greatly reduces the number of iterations of each training. In this way, it can not only track the dynamic changes of slow-changing parameters online, but also make real-time decisions according to the current fast-changing parameters, thus ensuring the generalization performance and real-time performance of the model. And unless the machine learning model trained offline is changed, there is no need to download the machine learning model to the online training system after the offline training is completed, but only the relevant parameters of the offline training need to be downloaded, thereby improving the efficiency of the entire training system. Compared with the above traditional offline training method, the online training of the present application only needs to learn the fast-changing parameters in the environmental parameters and track the slow-changing parameters, which makes the neural network structure in the online training method simpler than that in the offline training, has a faster real-time decision-making speed, and has a stronger adaptability to the environment.
[0220] Figure 8 It is a structural block diagram of a wireless resource allocation device provided by the embodiments of the present application. As Figure 8 shown, the wireless resource allocation device 800 may include: a first acquisition unit 801, a second acquisition unit 802, a third acquisition unit 803, an online training unit 804, and a resource allocation unit 805.
[0221] Specifically, the first acquisition unit 801 is used to acquire the large-scale channel gain at the current moment and the small-scale channel gain at the current moment in the dynamic environment;
[0222] The second acquisition unit 802 is used to acquire the small-scale channel gain target value; in some embodiments of the present application, the second acquisition unit 802 is specifically configured to: acquire the small-scale channel gain from the offline environment, and use the small-scale channel gain acquired from the offline environment as the small-scale channel gain target value.
[0223] The third acquisition unit 803 is used to acquire the parameters of the solver at the previous moment; the parameters include the user bandwidth, multiplier, and power neural network parameters;
[0224] The online training unit 804 is used to online train the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the small-scale channel gain target value, so as to update the parameters of the solver at the current moment;
[0225] In some embodiments of the present application, the online training unit 804 is specifically configured to: calculate the corresponding gradient according to the large-scale channel gain at the current moment; use the parameters of the solver at the previous moment as the initial value of the iterator at the current moment; perform N iterations on the iterator at the current moment according to the corresponding gradient and the small-scale channel gain target value; use the parameters obtained after N iterations of the iterator at the current moment as the parameters of the solver at the current moment.
[0226] In the embodiments of the present application, the formula of the iterator is expressed as follows:
[0227]
[0228]
[0229]
[0230] Where is the numerical implementation of the Lagrangian function at the t-th iteration, is the n-th realization of the achievable rate of the k-th user at the t-th iteration, N b is the batch data size, K is the number of users, is the effective bandwidth of the k-th user, θ k is a constant that can be obtained, E g is to take the expectation of the small-scale channel gain, ω(t) is the parameter of the power neural network at the t-th iteration, ω(t + 1) is the parameter of the power neural network at the (t + 1)-th iteration, and φ(t) is the learning rate at the t-th iteration, is the vector differential operator, W k is the bandwidth of the k-th user, W k (t) is the bandwidth of the k-th user at the t-th iteration, W k (t + 1) is the bandwidth of the k-th user at the (t + 1)-th iteration, λ kis the long-term constraint multiplier for the k-th user, λ k (t) is the long-term constraint multiplier for the k-th user at the t-th iteration, λ k (t + 1) is the long-term constraint multiplier for the k-th user at the (t + 1)-th iteration.
[0231] The resource allocation unit 805 is used to perform resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment. In some embodiments of the present application, the resource allocation unit is specifically used to: obtain the power of each user at the current moment according to the power neural network parameters among the parameters of the solver at the current moment and the small-scale channel gain at the current moment; determine the user bandwidth among the parameters of the solver at the current moment as the bandwidth of each user at the current moment; and perform resource allocation according to the bandwidth and power of each user at the current moment.
[0232] It should be noted that, in some embodiments of the present application, the model used by the iterator for the first online training is the iterator model pre-trained offline; wherein, the parameters in the iterator model pre-trained offline are used as the initial parameters of the solver when the iterator is first online trained.
[0233] It should be noted that the division of units in the embodiments of the present application is illustrative, merely a logical function division, and there may be other division methods in actual implementation. In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0234] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a processor-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0235] It should be noted here that the above device provided by the embodiments of the present invention can implement all the method steps implemented by the above method embodiments and can achieve the same technical effects. Therefore, the same parts and beneficial effects as those in the method embodiments will not be specifically described herein again.
[0236] Figure 9 It is a structural block diagram of a machine learning model training and inference device provided by an embodiment of the present application that adapts to dynamic changes in the communication environment. As Figure 9 described, the machine learning model training and inference device 900 may include: a first acquisition unit 901, a pre-training unit 902, a second acquisition unit 903, an online training unit 904, and an inference unit 905.
[0237] Specifically, the first acquisition unit 901 is used to acquire slow-changing parameters and fast-changing parameters in the offline environment.
[0238] The pre-training unit 902 is used to pre-train a preset iterator model by using unsupervised deep learning technology, slow-changing parameters, and fast-changing parameters in the offline environment
[0239] The second acquisition unit 903 is used to acquire slow-changing parameters at the current moment and fast-changing parameters at the current moment in the dynamic environment.
[0240] The online training unit 904 is used to online train the iterator according to the slow-changing parameters at the current moment in the dynamic environment, the parameters of the solver at the previous moment, and the fast-changing parameters in the offline environment to update the parameters of the solver at the current moment; among them, the model used when the iterator is first online trained is the iterator model pre-trained offline, and the parameters in the iterator model pre-trained offline are used as the initial parameters of the solver when the iterator is first online trained.
[0241] The inference unit 905 is used to perform real-time inference according to the parameters of the solver at the current moment and the fast-changing parameters at the current moment.
[0242] It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated units may be implemented in the form of hardware or in the form of software functional units.
[0243] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0244] It should be noted here that the above-mentioned device provided in the embodiment of the present invention can implement all the method steps implemented in the above-mentioned method embodiment and can achieve the same technical effect. The same parts and beneficial effects as those in the method embodiment will not be specifically described in this embodiment.
[0245] In an exemplary embodiment, a processor-readable storage medium is also provided. For example, the processor-readable storage medium in a network-side device stores a computer program, and the computer program is used to cause the processor to execute to complete the above method. For example, the storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0246] In an exemplary embodiment, a computer program product including a computer program is also provided. For example, when the computer program is executed by the processor of a network-side device, the above method can be completed.
[0247] Those skilled in the art will readily think of other implementation schemes of this application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application, and these variations, uses, or adaptations follow the general principles of this application and include common general knowledge or conventional technical means in the technical field not disclosed in this application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of this application are pointed out by the following claims.
[0248] It should be understood that this application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is only limited by the appended claims.
Claims
1. A method for wireless resource allocation, characterized in that, Including: Obtaining the large-scale channel gain at the current moment in a dynamic environment and the small-scale channel gain at the current moment; Obtaining the target value of the small-scale channel gain and obtaining the parameters of the solver at the previous moment; the parameters include user bandwidth, multiplier, and power neural network parameters; Online training the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment; Performing resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment.
2. The method according to claim 1, characterized in that, The obtaining of the target value of the small-scale channel gain includes: Obtaining the small-scale channel gain from an offline environment and using the small-scale channel gain obtained from the offline environment as the target value of the small-scale channel gain.
3. The method according to claim 1, characterized in that, The online training the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment includes: Calculating the corresponding gradient according to the large-scale channel gain at the current moment; Using the parameters of the solver at the previous moment as the initial value of the iterator at the current moment; Performing N iterations on the iterator at the current moment according to the corresponding gradient and the target value of the small-scale channel gain; Using the parameters obtained after the N iterations of the iterator at the current moment as the parameters of the solver at the current moment.
4. The method according to claim 3, characterized in that, The formula of the iterator is expressed as follows: Among them is the numerical implementation of the Lagrangian function at the t-th iteration, is the n-th realization of the achievable rate of the k-th user at the t-th iteration, N b is the batch data size, K is the number of the users, is the effective bandwidth of the k-th user, θ k is a constant that can be obtained, E g is to take the expectation of the small-scale channel gain, ω(t) is the parameter of the power neural network at the t-th iteration, ω(t + 1) is the parameter of the power neural network at the (t + 1)-th iteration, and φ(t) is the learning rate at the t-th iteration, is the vector differential operator, W k is the broadband of the k-th user, W k W(t) is the broadband of the k-th user at the t-th iteration, W k W(t + 1) is the broadband of the k-th user at the (t + 1)-th iteration, λ k is the long-term constraint multiplier of the k-th user, λ k λ(t) is the long-term constraint multiplier of the k-th user at the t-th iteration, λ k λ(t + 1) is the long-term constraint multiplier of the k-th user at the (t + 1)-th iteration.
5. The method according to claim 1, characterized in that, The performing resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment includes: Obtaining the power of each user at the current moment according to the power neural network parameters among the parameters of the solver at the current moment and the small-scale channel gain at the current moment; Determining the user bandwidth among the parameters of the solver at the current moment as the bandwidth of each user at the current moment; Performing resource allocation according to the bandwidth and power of each user at the current moment.
6. The method according to claim 1, characterized in that, The model used when the iterator performs online training for the first time is an iterator model pre-trained offline; wherein, the parameters in the iterator model pre-trained offline are used as the initial parameters of the solver when the iterator is performing online training for the first time.
7. A network-side device, characterized in that, Including a memory, a transceiver, and a processor; wherein, The memory is used to store computer programs; the transceiver is used to transmit and receive data under the control of the processor; the processor is used to read the computer programs in the memory and perform the following operations: Obtaining the large-scale channel gain at the current moment in a dynamic environment and the small-scale channel gain at the current moment; Obtaining the target value of the small-scale channel gain and obtaining the parameters of the solver at the previous moment; the parameters include user bandwidth, multiplier, and power neural network parameters; Online training the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the target value of the small-scale channel gain to update the parameters of the solver at the current moment; Performing resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment.
8. The network-side device according to claim 7, characterized in that, The obtaining of the small-scale channel gain target value includes: Obtaining the small-scale channel gain from an offline environment, and using the small-scale channel gain obtained from the offline environment as the small-scale channel gain target value.
9. The network-side device according to claim 7, characterized in that, The online training of the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the small-scale channel gain target value to update the parameters of the solver at the current moment includes: Calculating the corresponding gradient according to the large-scale channel gain at the current moment; Using the parameters of the solver at the previous moment as the initial value of the iterator at the current moment; Performing N iterations on the iterator at the current moment according to the corresponding gradient and the small-scale channel gain target value; Using the parameters obtained after the N iterations of the iterator at the current moment as the parameters of the solver at the current moment.
10. The network-side device according to claim 9, wherein, The formula of the iterator is expressed as follows: Among them is the numerical implementation of the Lagrangian function at the t-th iteration, is the n-th implementation of the achievable rate of the k-th user at the t-th iteration, N b is the batch data size, K is the number of the users, is the effective bandwidth of the k-th user, θ k is a constant that can be obtained, E g is to take the expectation of the small-scale channel gain, ω(t) is the parameter of the power neural network at the t-th iteration, ω(t + 1) is the parameter of the power neural network at the (t + 1)-th iteration, and φ(t) is the learning rate at the t-th iteration, is the vector differential operator, W k is the broadband of the k-th user, W k W(t) is the broadband of the k-th user at the t-th iteration, W k W(t + 1) is the broadband of the k-th user at the (t + 1)-th iteration, λ k is the long-term constraint multiplier of the k-th user, λ k λ(t) is the long-term constraint multiplier of the k-th user at the t-th iteration, λ k λ(t + 1) is the long-term constraint multiplier of the k-th user at the (t + 1)-th iteration.
11. The network-side device according to claim 7, wherein, The resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment includes: Obtaining the power of each user at the current moment according to the power neural network parameters among the parameters of the solver at the current moment and the small-scale channel gain at the current moment; Determining the user bandwidth among the parameters of the solver at the current moment as the bandwidth of each user at the current moment; Performing resource allocation according to the bandwidth and power of each user at the current moment.
12. The network-side device according to claim 7, wherein, The model used when the iterator is first online trained is an iterator model pre-trained offline; wherein, the parameters in the iterator model pre-trained offline are used as the initial parameters of the solver when the iterator is first online trained.
13. A wireless resource allocation device, wherein, It includes: A first obtaining unit, configured to obtain the large-scale channel gain at the current moment and the small-scale channel gain at the current moment in a dynamic environment; A second obtaining unit, configured to obtain the small-scale channel gain target value; A third obtaining unit, configured to obtain the parameters of the solver at the previous moment; the parameters include user bandwidth, multiplier, and power neural network parameters; An online training unit, configured to online train the iterator according to the large-scale channel gain at the current moment, the parameters of the solver at the previous moment, and the small-scale channel gain target value to update the parameters of the solver at the current moment; A resource allocation unit, configured to perform resource allocation according to the parameters of the solver at the current moment and the small-scale channel gain at the current moment.
14. A processor-readable storage medium, wherein, The processor-readable storage medium stores a computer program, and the computer program is used to cause the processor to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Computing resource allocation and task unloading method for edge computing of super-dense network
CN110798849A
Physical layer security resource allocation method in ICV network
CN112153744A