Downlink power adjustment method and device, electronic equipment and storage medium
By utilizing neural networks and reinforcement learning algorithms to predict channel quality in short time units within a wireless communication system, fine-grained adjustment of downlink power is achieved. This solves the problem of timely acquisition and dynamic allocation of channel quality in high-speed mobile scenarios, thereby improving the user experience.
Patent Information
- Application Number
- CN202511467017.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-02-13
AI Technical Summary
In wireless communication systems, existing technologies struggle to achieve timely acquisition of downlink channel quality and dynamic power allocation in high-speed mobile scenarios, resulting in a poor user experience.
By acquiring the channel quality reported by the terminal, using neural network models and reinforcement learning algorithms, the channel quality in shorter time units is predicted, and the downlink power is adjusted according to the prediction results to achieve fine-grained power allocation.
It improves the flexibility and accuracy of downlink power allocation, enhancing the user's service experience, especially in high-speed mobile scenarios.
Smart Images

Figure CN121531441A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication, in particular to a downlink power adjustment method and device, electronic equipment and storage medium. BACKGROUND
[0002] In a wireless communication system, through downlink static power allocation by a network side device (for example, a base station), the power on all RB (Resource Block) resources scheduled by the user in the entire scheduling process is the same and unchangeable. The static power allocation scheme in the related art is difficult to meet the complex and continuously changing network ecology and user experience requirements.
[0003] Further, for example, when the terminal is in a high-speed moving scenario such as a high-speed train, the contradiction is more prominent. When the user is moving at high speed, the level of the downlink channel changes quickly, and the period of the terminal reporting CSI (Channel State Information) in the related art is long, and the feedback is not timely, which will also make the network side device unable to obtain the channel quality in time, so as to unable to dynamically adjust the downlink power allocation.
[0004] How to optimize the related technology and improve the user's service experience has become a technical problem to be solved. SUMMARY
[0005] An embodiment of the present application provides a downlink power adjustment method, device, electronic equipment and storage medium, so as to improve the timeliness of the network side device obtaining the downlink channel information, improve the flexibility of the downlink power allocation, and improve the user's service experience.
[0006] To solve the above technical problems, an embodiment of the present application is implemented as follows: In a first aspect, an embodiment of the present application provides a downlink power adjustment method, which comprises: obtaining a first channel quality reported by a terminal; processing the first channel quality to obtain a second channel quality corresponding to at least one first time unit; adjusting a second downlink power of a target time unit according to the second channel quality corresponding to the first time unit and a first downlink power, to obtain an adjusted third downlink power; Wherein, the length of the first time unit is less than the minimum reporting period of the first channel quality, and the target time unit is at least one first time unit whose power is to be adjusted.
[0007] In a second aspect, another embodiment of the present application provides a downlink power adjustment device, which comprises: obtaining a first channel quality reported by a terminal; predicting a second channel quality corresponding to at least one first time unit based on the first channel quality; determining a second downlink power of a target time unit based on the second channel quality corresponding to the first time unit and the first downlink power, to obtain a third downlink power after adjustment; wherein a length of the first time unit is less than a minimum reporting period of the first channel quality, and the target time unit is at least one first time unit whose power is to be adjusted.
[0008] In a third aspect, a further embodiment of the present specification provides an electronic device, characterized in that the electronic device comprises a memory and a processor, and the memory has stored computer executable instructions, and the computer executable instructions, when executed on the processor, can implement the steps of the downlink power adjustment method of the first aspect.
[0009] In a fourth aspect, a further embodiment of the present specification provides a computer readable storage medium for storing computer executable instructions, and the computer executable instructions, when executed on a processor, can implement the steps of the downlink power adjustment method of the first aspect.
[0010] In a fifth aspect, a further embodiment of the present specification provides a computer program product, comprising a processing program, and the processing program, when executed on a processor, can implement the steps of the downlink power adjustment method of the first aspect.
[0011] The embodiments of the present disclosure obtain a first channel quality reported by a terminal; process the first channel quality to obtain a second channel quality corresponding to at least one first time unit; and adjust a second downlink power of a target time unit based on the second channel quality corresponding to the first time unit and the first downlink power, to obtain a third downlink power after adjustment; wherein a length of the first time unit is less than a minimum reporting period of the first channel quality, and the target time unit is at least one first time unit whose power is to be adjusted. The embodiments of the present disclosure have the following advantages or beneficial effects: the channel quality of the target time unit whose power is to be adjusted can be accurately predicted, the downlink power can be finely adjusted, the flexibility of downlink power allocation is greatly improved, and the user's service experience can be improved.
[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure.
[0013] Other features and advantages of the present disclosure will be described in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to make one or more embodiments of the present disclosure better understood, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present disclosure, and for those skilled in the art, other drawings can be obtained without creative labor.
[0015] Figure 1 A flow chart of a downlink power adjustment method provided by one embodiment of the present disclosure; Figure 2 A flow chart of another downlink power adjustment method provided by one embodiment of the present disclosure; Figure 3 A flow chart of a Q value table training method provided by one embodiment of the present disclosure; Figure 4 A flow chart of another downlink power adjustment method provided by one embodiment of the present disclosure; Figure 5 A flow chart of a model training method provided by one embodiment of the present disclosure; Figure 6 An architecture schematic diagram of a model training provided by one embodiment of the present disclosure; Figure 7 A flow chart of another downlink power adjustment method provided by one embodiment of the present disclosure; Figure 8 A schematic diagram of a downlink power adjustment device provided by one embodiment of the present disclosure; Figure 9 A schematic diagram of another downlink power adjustment device provided by one embodiment of the present disclosure; Figure 10 A hardware structure schematic diagram of an electronic device provided by one embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] In order to make one or more embodiments of the present disclosure better understood, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present disclosure, and for those skilled in the art, other drawings can be obtained without creative labor.
[0017] The terminal described in the embodiments of the present application can also be referred to as a user equipment (UE), and can be a terminal-side device such as a mobile phone, a tablet personal computer, a laptop computer, a notebook computer, a personal digital assistant (PDA), a palm computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile Internet device (MID), an augmented reality (AR) device, a virtual reality (VR) device, and the like. The network-side device described in the embodiments of the present application can include an access network device, which can also be referred to as a radio access network (RAN) device, a radio access network function or a radio access network unit. The access network device can include a base station, which can also be referred to as a node B (NB), an evolved node B (eNB), a next generation node B (gNB), a new radio node B (NR NodeB), an access point or some other appropriate terminology in the art, as long as the same technical effects are achieved. The base station is not limited to a specific technical term, and it should be noted that only the base station in the NR system is taken as an example for introduction in the embodiments of the present application, and the specific type of the base station is not limited.
[0018] It is worth noting that the technology described in the embodiments of the present application is not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA) or other systems. The terms "system" and "network" in the embodiments of the present application are often used interchangeably, and the described technology can be used in the above-mentioned systems and radio technologies, as well as other systems and radio technologies. The following description describes a New Radio (NR) system for example purposes, and NR terminology is used in most of the following description, but these technologies can also be applied to systems other than NR systems, such as 6th Generation (6G) communication systems. th
[0019] An embodiment of a downlink power adjustment method provided by the present specification is described with reference to Figure 1 which shows a flowchart of a downlink power adjustment method provided by the present embodiment, which can be run in a network side device (such as a base station), for example, an intelligent single board can be deployed locally in the base station, and the downlink power adjustment method is run through the software on the intelligent single board, the downlink power adjustment provided by the present embodiment specifically includes the following steps.
[0020] In step S101, a first channel quality reported by a terminal is obtained.
[0021] In some embodiments, the network side device can receive the first channel quality reported by the terminal, and the channel quality can be CSI (Channel State Information), for example, the channel quality can be CQI (Channel Quality Indication).
[0022] A terminal can perform periodic measurement on a reference signal sent by a network side device (e.g., a base station), and report a first channel quality to the network side device, where the reference signal is, for example, a CSI-RS (Channel State Information Reference Signal) in NR (New Radio).
[0023] In some embodiments, a period of reporting the first channel quality by the terminal (e.g., a second time unit) can meet a reporting period of the channel quality specified in a standard protocol, for example, in NR, the period can be configured as 40 ms, 80 ms, or 160 ms, in other words, the terminal can report the first channel quality with a granularity of a second time unit with a length of 40 ms, 80 ms, or 160 ms.
[0024] In step S102, the first channel quality is processed to obtain a second channel quality corresponding to at least one first time unit.
[0025] The length of the first time unit is less than a minimum reporting period of the first channel quality.
[0026] For example, the first time unit can be a time slot, for example, in NR 15Khz SCS (Subcarrier Spacing), the first time unit can be 1 ms, and in NR 30Khz SCS, the first time unit can be 0.5 ms. Only the length of the first time unit needs to be less than the minimum reporting period of the channel quality specified in the standard protocol. Of course, the first time unit can also be multiple time slots, for example, it can be 2 time slots.
[0027] In some embodiments, the neural network model can be trained, and the trained neural network model can be used to process (e.g., interpolate and predict) the channel quality, for example, when the terminal reports the first channel quality with a period of 40 ms, the trained neural network model can be used to predict the channel quality, thereby obtaining the second channel quality corresponding to multiple first time units (e.g., with a length of 1 ms) between two reports of the first channel quality by the terminal.
[0028] It is understandable that channel quality can also be rank indicator (RI), precoding matrix indicator (PMI), or any combination of CQI, RI, and PMI. In some embodiments, after obtaining the second channel quality for a period shorter than the minimum reporting period (i.e., the first time unit, e.g., each time slot) specified in the standard protocol, in addition to determining the power adjustment action for the next first time unit, it can also be used to optimize terminal scheduling. For example, adaptive modulation and coding (AMC) can be optimized based on the second channel quality corresponding to each first time unit to further optimize the user experience.
[0029] In step S103, the second downlink power of the target time unit is adjusted according to the second channel quality and the first downlink power corresponding to the first time unit to obtain the adjusted third downlink power.
[0030] In some embodiments, the target time unit can be at least one future first time unit, i.e. at least one first time unit of power to be adjusted. Since the channel quality is closely related to the downlink power allocated by the network-side device, the downlink power of the target time unit can be adjusted according to the second channel quality and the first downlink power corresponding to the first time unit, thereby improving the accuracy of power adjustment. It can be understood that the network-side device can obtain the first downlink power allocated in the same first time unit corresponding to the second channel quality, or in other words, the network-side device can obtain the second channel quality and the first downlink power corresponding to each historical first time unit, and adjust the downlink power of at least one future target time unit of power to be adjusted according to this information.
[0031] For example, when the first time unit is a single time slot, the corresponding first downlink power is the downlink power allocated by the network-side device in that first time unit. When the first time unit is multiple time slots, the first downlink power can be the average of the downlink power allocated by the network-side device in multiple time slots.
[0032] Figure 2 A flowchart of another downlink power adjustment method provided in one embodiment of this specification is shown below. Figure 2 As shown, step S103 may include the following steps.
[0033] In step S1031, the power adjustment value of the target time unit is determined based on the second channel quality and the first downlink power corresponding to the first time unit.
[0034] In some embodiments, a reinforcement learning algorithm can be employed to determine the power adjustment value of the target time unit according to the second channel quality and the first downlink power corresponding to the first time unit, and of course other suitable algorithms can also be employed, which are not limited in the present application.
[0035] For example, a reinforcement learning algorithm can be employed to determine the power adjustment value of the N+k target time unit of the power to be adjusted according to the second channel quality and the first downlink power corresponding to the N first time unit, where k can be a positive integer greater than or equal to 1. Of course, the power adjustment action of the N+k target time unit of the power to be adjusted can also be determined according to the second channel quality and the first downlink power corresponding to the N to N+k-1 first time unit respectively.
[0036] In some embodiments, the power adjustment action is a power adjustment value corresponding to at least one preset bandwidth respectively; the reinforcement learning algorithm is a Q-learning learning algorithm; the state space of the Q-learning learning algorithm is a two-dimensional state space corresponding to at least one preset bandwidth respectively, and the two-dimensional state space dimension includes the channel quality reported by the terminal and the corresponding downlink power; the action space of the Q-learning learning algorithm is a power adjustment value corresponding to at least one preset bandwidth respectively; and the reward function of the Q-learning learning algorithm is a weighted average value of the channel quality reported by the terminal corresponding to at least one preset bandwidth respectively. The preset bandwidth includes any of the following: A first preset number of resource blocks (RBs); A second preset number of sub-bands.
[0037] In some embodiments, the second channel quality and the first downlink power can be taken as the current state, and the power adjustment value corresponding to the current state can be obtained from the pre-trained Q value table. It can be understood that, since the state space includes the second channel quality and the first downlink power corresponding to the preset bandwidth, the power adjustment value can also be the power adjustment value corresponding to the preset bandwidth, and the power adjustment values of different preset bandwidths can be the same or different, that is, the power adjustment value can be a multi-dimensional power adjustment vector including power adjustment values corresponding to multiple preset bandwidths.
[0038] For example, the state space of the wireless network system is defined as Each state is composed of the second channel quality and the first downlink power corresponding to a certain first time unit by at least one user. are vectors of second channel quality and first downlink power on the full bandwidth respectively, each element of the vector of second channel quality represents the second channel quality corresponding to a preset bandwidth of frequency domain resource, and each element of the vector of first downlink power represents the first downlink power corresponding to a preset bandwidth of frequency domain resource. For example, the full bandwidth is 100 RBs, and the preset bandwidth is 4 RBs, and each is a 100 / 4=25-dimensional vector. The size of the state space is a two-dimensional state space composed of all possible values of the second channel quality and the first downlink power.
[0039] For example, the power adjustment action space A is defined as Each action A represents a vector composed of power adjustment values of each preset bandwidth of frequency resource on the full bandwidth. Each power adjustment value in the vector can include increasing a unit of power (+1), decreasing a unit of power (-1), or keeping the power unchanged (0), and of course can include adjusting multiple units of power (for example, +2, -2, etc.). For example, the full bandwidth is 100 RBs, and the preset bandwidth is 4 RBs, and each action is a 100 / 4=25-dimensional vector, and the power adjustment value can be, for example, A1=(+1, +1, +1, -1, -1, 0, 0, …), each element of the vector represents a power adjustment value corresponding to a preset bandwidth of frequency domain resource, and all possible values of the vector constitute the above-mentioned action space.
[0040] In some embodiments, the reward function of the Q-learning learning algorithm can be as shown in Formula One: (Formula One) Wherein, is the total number of preset bandwidths of frequency domain resources within the full bandwidth, is the scheduling situation in the downlink channel, if the preset bandwidth of resource is scheduled, then , otherwise , is the second channel quality corresponding to the kth preset bandwidth of the u user in the first time unit.
[0041] In the embodiments of the present application, the power allocated by the downlink channel on the preset bandwidth k The relationship between the downlink channel quality allocated by the downlink channel on the preset bandwidth k and the CQI reported or predicted by the user is no longer a simple linear relationship, but a power-constrained allocation model with the overall downlink channel quality of the wireless system as the target, as shown in Formula Two: (Formula Two) Wherein, is the power allocated on the kth preset bandwidth, that is, the third downlink power adjusted on the kth preset bandwidth, and Pmax is the total transmit power.Constraints of characterizing power adjustment.
[0042] In some embodiments, by employing the reinforcement learning algorithm described above, the power adjustment value of the target time unit t+1 (e.g., the t+1th time slot) can be determined according to the second channel quality of the first time unit t (e.g., the tth time slot) and the corresponding first downlink power.
[0043] In step S1032, the second downlink power of the target time unit is adjusted according to the power adjustment value to obtain the adjusted third downlink power.
[0044] It can be understood that the power adjustment action can be a power adjustment value for a preset bandwidth, and the downlink power can be adjusted according to the power adjustment value in the frequency domain with the preset bandwidth and the first time unit as the granularity in the time domain. For example, the default RE (Resource element) power of a wireless cell is related to the device transmission power, cell bandwidth, and subcarrier spacing. Taking a cell with a 120W AAU, 100M bandwidth, and 30KHz subcarrier spacing as an example, the RE reference power is 15.6dBm, and the PDSCH (Physical Downlink Shared Channel) second downlink power of the t time period before adjustment and the t+1 time period after adjustment can be shown in Table 1 and Table 2 respectively. The units in the table are all in dBm.
[0045] Table 1
[0046] Table 2
[0047] As shown in Table 1 and Table 2, different rows of the table represent different frequency domain resources (which can correspond to different preset bandwidths), different columns represent different time domain resources (which can correspond to different symbols within the same first time unit), and cells (1, 1) to (2, 7) represent that the downlink power corresponding to service A of user 1 is adjusted from 15.6dBm to 15.2dBm, cells (3, 1) to (4, 7) represent that the downlink power corresponding to service B of user 1 is adjusted from 15.6dBm to 16.8dBm, and cells (5, 1) to (10, 7) represent that the downlink power corresponding to user 2 is adjusted from 15.6dBm to 14.2dBm.
[0048] By adopting the technical scheme, the granularity of adjusting power can be optimized to the preset bandwidth and the first time unit granularity, so that the downlink power can be adjusted according to the preset bandwidth and the first time unit as the granularity in a fine manner without exceeding the maximum transmission power, the flexibility of downlink power distribution is greatly improved, the best power adjustment strategy can be determined in time and accurately for users including high-speed moving users, and the service experience of the users can be improved.
[0049] In some embodiments, the user weight and / or the service weight can also be considered for further differentiated power adjustment, and the weight of the reward function includes at least one of the following: the user weight corresponding to the terminal; The service weight corresponding to different service types.
[0050] For example, assuming that the set of users scheduled in the network side device (for example, a base station) is The user category u weight corresponding to the terminal can be Assuming that the service type set of the user u is V The service weight corresponding to the service type v can be The weight of the reward function can be .
[0051] The reward function considering the weight can be as shown in the following formula three: (Formula three) Compared with formula one, formula three adds the weight coefficient , which is the preset weight of the u th user and the v th service.
[0052] Correspondingly, the allocation model considering the power constraint and taking the downlink channel quality of the whole wireless system as the target is as shown in the following formula four: (Formula four) Compared with formula two, formula four adds the weight coefficient .
[0053] In some possible implementation manners, the correspondence between each user category and the user weight can be predefined as shown in the example of table three, and the correspondence between each service type and the service weight can be predefined as shown in the example of table four, and The following table three and table four can be queried respectively.
[0054] Table three
[0055] Table four
[0056] By adopting the technical scheme, when the downlink power is allocated, not only the influence of the channel quality is considered, but also the influence of the service type and / or the user category is further considered, so that the differentiated service experience of the user of the corresponding service type and / or user category is further improved on the basis of improving the flexibility of the downlink power allocation.
[0057] Figure 3 A flowchart of a Q value table training method provided for an embodiment of the present specification is shown in Figure 3 The Q value table corresponding to the Q-learning learning algorithm is obtained by the following method.
[0058] Step S201 is step 1, and the initial value of the Q value table is a preset value.
[0059] In the embodiment of the present application, the Q value table is used to record the values corresponding to different power adjustment values of the wireless system in different states, and the Q value table is a two-dimensional table composed of a state space and a power adjustment value space, and the size of the table is the product of the size of the state space and the size of the power adjustment value space.
[0060] In some embodiments, the Q values of all state-action pairs can be initialized to 0 or a random positive value less than a second threshold.
[0061] Step S202 is step 2, and a first adjustment value is determined from the power adjustment value space according to a preset selection policy based on a first state corresponding to a current first time unit, and the Q value table corresponding to the first state and the first adjustment value is updated, and the first state includes the channel quality and the downlink power corresponding to the current first time unit.
[0062] In some embodiments, the update formula of the Q value is shown in formula five.
[0063] (Formula five) wherein, represents a preset learning rate for controlling the balance between the new Q value and the old Q value, represents a preset discount factor for controlling the weight of the future reward, is the maximum Q value in all possible power adjustment value spaces in the new state , represents the maximum expectation of the future reward, t represents the current first time unit, and t+1 represents the next first time unit.
[0064] The update formula of the Q value can be selected according to the current state in the Q value table to perform the action that maximizes the Q value, and after the action is performed, a reward is obtained, and the Q value table can be updated, and the update target of the Q value table is to make the Q value continuously approach the actual reward that should be obtained.
[0065] In some embodiments, the preset selection strategy can be Strategy, specifically, the action with the highest current Q value can be selected with a probability of 1- , and an action can be randomly selected with a probability of 1- , where is a preset value within (0, 1).
[0066] Step S203 is step 3, and the next first time unit is taken as the current time unit, and step 2 is repeated until the Q value table converges, and a trained Q value table is obtained.
[0067] Through iterative training of large-scale data in the live network, a converged Q value table can be obtained, and in some embodiments, the converged Q value table can be deployed in a corresponding network side device (for example, a base station), for example, the converged Q value table can be deployed in an intelligent single board of the base station. The condition for the convergence of the Q value table can be that the average number of updates of each table entry of the Q value table reaches a preset number threshold, or the change of the table entry of the Q value table is less than a preset change threshold, which is not limited in the present application.
[0068] With the above technical solution, each network side device (for example, a base station) can update the corresponding Q value table according to its own wireless environment and traffic model through iterative training of large-scale data of actual terminals in the live network, and different network side devices (for example, base stations) can learn different Q value tables, thereby being able to realize the capability of individually intelligent adjusting downlink power.
[0069] Figure 4 The flowchart of another downlink power adjustment method provided by an embodiment of the present specification is shown in Figure 4 , and step S102 can include the following steps.
[0070] In step S1021, the first channel quality and the corresponding fourth downlink power are inferred through the pre-trained neural network model to obtain the second channel quality corresponding to at least one first time unit output by the neural network model.
[0071] The loss function of the neural network model includes a first loss function, and the first loss function represents the correlation between the predicted second channel quality and the corresponding second downlink power.
[0072] In some embodiments, the first channel quality, the second channel quality, and the second downlink power correspond to a plurality of preset bandwidths within the full bandwidth; wherein the preset bandwidths include any one of the following: a first preset number of RBs (Resource Blocks); a second preset number of subbands.
[0073] In an example, the preset bandwidth can be 1 RB, i.e., the first channel quality corresponding to each RB is fed back by the terminal, or the preset bandwidth can be multiple RBs (e.g., 5 RBs), i.e., the first channel quality corresponding to each 5 RBs is fed back by the terminal. The preset bandwidth can be one subband, i.e., the first channel quality corresponding to each subband is fed back by the terminal, or the preset bandwidth can be multiple subbands (e.g., 2 subbands), i.e., the first channel quality corresponding to each 2 subbands is fed back by the terminal. The concept and definition of the subband can be referred to the related protocol standard.
[0074] The first channel quality and the corresponding second downlink power reported by the terminal can form an N*2 matrix, where N is the number of preset bandwidths in the full bandwidth. For example, the full bandwidth includes 100 RBs, and the preset bandwidth is 4 RBs, and the terminal reports the first channel quality in each second time unit corresponding to a 25*2 matrix.
[0075] In some embodiments, the matrix of the first channel quality and the fourth downlink power corresponding to the multiple second time units in which the terminal reports the first channel quality can be input into the pre-trained neural network model, so as to obtain the second channel quality corresponding to at least one first time unit in the future. It can be understood that the number of the above terminals can be one or multiple. That is, different terminals can feed back the first channel quality corresponding to different preset bandwidths.
[0076] In some embodiments, the loss function of the pre-trained neural network model can include a first loss function as shown in Formula Six.
[0077] (Formula Six) wherein, is the first loss function, and denote the partial derivatives of the predicted second channel quality and the corresponding first downlink power with respect to time t, denotes the Pearson correlation coefficient.
[0078] In the training of the above pre-trained neural network model, the second channel quality and the corresponding first downlink power of multiple different first time units predicted by the neural network model to be trained are regarded as a binary function of time t, and the partial derivatives of the predicted second channel quality and the corresponding first downlink power with respect to time t are obtained. By taking the value of the above first loss function tending to a preset value (e.g., 0) as the optimization goal, the neural network model to be trained is optimized by back propagation, which can make the physical correlation of the second channel quality and the first downlink power consistent, so that the prediction of the second channel quality is more in line with the physical law.
[0079] By using the technical scheme, the channel quality of at least one target time unit in the future can be accurately predicted through the first channel quality reported by the terminal and the pre-trained neural network model, so as to perform downlink power adjustment according to the predicted second channel quality of the target time unit, and improve user experience.
[0080] In some embodiments, the loss function of the pre-trained neural network model further includes at least one of a second loss function and a third loss function, the second loss function representing a mean square error of the second channel quality and the corresponding first channel quality, and the third loss function representing a distribution error of the second channel quality and the corresponding first channel quality.
[0081] The second loss function can be as shown in Formula Seven: (Formula Seven) Wherein, is the second loss function, is the i th first channel quality, is the i th second channel quality, is the number of first time units.
[0082] The third loss function can be as shown in Formula Eight: (Formula Eight) Wherein, is the third loss function, is the i th first channel quality, is the i th second channel quality, is the number of first time units, is a softmax function, used to convert the sequence of first channel quality or second channel quality into a distribution to calculate the distribution divergence.
[0083] In some embodiments, the first loss function, and at least one of the second loss function and the third loss function can be weighted to obtain a final loss function.
[0084] For example, when the loss function includes the first loss function, the second loss function, and the third loss function, the final loss function can be determined by Formula Nine as follows.
[0085] (Formula Nine) Wherein, is the final loss function, , , respectively are preset weights of the corresponding loss functions.
[0086] By adopting the technical scheme, the distribution consistency of the predicted second channel quality and the first channel quality and the first downlink power obtained after the neural network model to be trained is optimized based on the final loss function can be improved, and the second channel quality corresponding to the target time unit can be accurately predicted even if the user is in a high-speed mobile state, thereby further improving the accuracy of the trained neural network model.
[0087] In some embodiments, the neural network model to be trained can be a long short-term memory neural network, for example, the neural network model to be trained can be an LSTM (Long Short-Term Memory), and of course, other suitable neural network models can also be used, which are not limited in the embodiments of the present application. Figure 5 A flowchart of a model training method provided by an embodiment of the present specification is shown in Figure 6 A schematic diagram of a model training architecture provided by an embodiment of the present specification is shown in Figure 5 As shown, the neural network model is trained by the following method.
[0088] Step S301 is step 1, obtaining the third channel quality reported by the training terminal in a plurality of third time units and the fifth downlink power corresponding to the third time unit.
[0089] It can be understood that the network side device can obtain the fifth downlink power allocated by the same third time unit associated with each third channel quality, or in other words, the network side device can obtain the third channel quality and the fifth downlink power corresponding to each third time unit, and train the neural network model according to these information.
[0090] In some embodiments, the length of the third time unit can be the same as or different from the length of the first time unit, and the third time unit can be a time slot, of course, the third time unit can also be a plurality of time slots, for example, it can be 2 time slots, and the third downlink power can be the average of the downlink power allocated to the training terminal. In some embodiments, the length of the third time unit can also be the same as the length of the second time unit, which is not limited in the present application.
[0091] In some embodiments, the training terminal can be a terminal customized for collecting training data, and of course, the training terminal can also be a communication device simulating the terminal to measure and feedback, and the training terminal can measure the reference signal sent by the network side device (for example, a base station) in different wireless scenarios by traversing different wireless scenarios, and report the corresponding third channel quality according to the third time unit.
[0092] It can be understood that when the period actually reported by the terminal is different from the length of the third time unit, the training data of the third time unit can be converted. For example, the third time unit is 1 ms, and the training terminal can only perform channel quality feedback according to a period of 4 ms. In some embodiments, the channel quality feedback by the training terminal can be linearly interpolated to obtain the third channel quality and the corresponding third downlink power corresponding to each third time unit, for example, the CQI feedback by the training terminal in the previous and next two 4 ms is 10 and 14 respectively, and linear interpolation can be performed to obtain the CQI corresponding to the five third time units (1 ms) as 10, 11, 12, 13, and 14.
[0093] It can be understood that when the training terminal feeds back the third channel quality, the feedback frequency domain granularity can be different from the preset bandwidth, and the third channel quality corresponding to multiple preset bandwidths can be obtained based on an equivalent conversion method. For example, the preset bandwidth is 1 RB, and the terminal feeds back the third channel quality according to the sub-band, and each sub-band includes 4 RBs, and it can be determined that the channel quality of the 4 RBs corresponding to the sub-band is the third channel quality. For another example, the preset bandwidth is 4 RBs, and the terminal feeds back the third channel quality according to the sub-band, and each sub-band includes 2 RBs, and the average of the channel quality corresponding to the 4 RBs of the preset bandwidth can be taken as the third channel quality corresponding to the preset bandwidth. Similarly, when the network side device (such as a base station) allocates the downlink power, the granularity of the allocation can also be different from the preset bandwidth, and the downlink power corresponding to each preset bandwidth can be determined based on a similar method, which will not be described here.
[0094] The third channel quality and the corresponding fifth downlink power reported by the training terminal in each third time unit can form an N*2 matrix, where N is the number of preset bandwidths in the full bandwidth. For example, the full bandwidth includes 100 RBs, and the preset bandwidth is 4, so each third time unit corresponds to a 25*2 matrix. In the embodiments of the present application, the matrices corresponding to multiple third time units can be used as training data to train the model. It can be understood that the third channel quality reported in each third time unit can also be used as a label to predict the fourth channel quality in the current third time unit.
[0095] Step S302 is step 2, and the multiple third channel qualities and fifth downlink powers except the last third channel quality and fifth downlink power in the preset sliding window are used as training data to train the long short-term memory neural network to be trained.
[0096] Step S303 is step 3, obtaining the fourth channel quality predicted by the long short-term memory neural network to be trained according to the training data, determining the result of the loss function according to the last third channel quality, the fourth channel quality, and the corresponding fifth downlink power of the preset sliding window, and performing back propagation optimization on the long short-term memory neural network to be trained according to the result of the loss function.
[0097] For example, as shown in Figure 6 , the preset sliding window length is , the training data input to the long short-term memory neural network to be trained at a certain time can be -1 third channel quality and the fifth downlink power corresponding matrix (i.e. multiple third channel qualities and fifth downlink powers except the last third channel quality and fifth downlink power), through Figure 5 , the output of the LSTM network hidden layer is input into the full connection layer to determine the predicted fourth channel quality output by the long short-term memory neural network to be trained, and then the predicted fourth channel quality output by the long short-term memory neural network to be trained and the last (i.e. the third channel quality and the corresponding fifth downlink power in the preset sliding window are used to calculate the result of the loss function, and the long short-term memory neural network to be trained is optimized by back propagation.
[0098] Step S304 is step 4, sliding the preset sliding window backward by the third preset number of third time units, repeating steps 2 and 3 until the result of the loss function meets the preset condition, and obtaining the trained long short-term memory neural network.
[0099] For example, the preset sliding window can be slid backward by one third time unit, and steps 2 and 3 can be repeated until the result of the loss function meets the preset condition, for example, the preset condition can be that the result of the loss function is less than a preset first threshold, and the trained long short-term memory neural network is obtained.
[0100] By using the above technical solution, the neural network model can be trained to accurately predict the channel quality of at least one third time unit to be predicted in the future, and the downlink power can be adjusted according to the predicted channel quality, thereby improving the user experience.
[0101] Figure 7 For another downlink power adjustment method provided by an embodiment of the present specification, a flowchart is shown in Figure 7 , the method is executed by an intelligent single board deployed on a network side device (such as a base station), an intelligent agent can be deployed on the intelligent single board, and the method can include the following steps.
[0102] Step 1, obtaining CQI reported by at least one terminal.
[0103] Step 2, performing CQI interpolation prediction, performing prediction on CQI in a first time unit granularity in a CQI reporting period to obtain CQI corresponding to each first time unit. In some embodiments, a trained neural network model can be deployed on an intelligent single board, and prediction can be performed by the trained neural network model. The trained neural network model can also be further iteratively updated by the actual CQI reported by the at least one terminal.
[0104] Step 3, identifying user categories and / or service types to determine differentiated weights of downlink power adjustment.
[0105] Step 4, determining a power adjustment value of a next first time unit according to the CQI corresponding to the current first time unit and the corresponding downlink power, so as to optimize the overall downlink channel quality. For example, the power adjustment value of the next time slot can be determined according to the state of the current time slot wireless system.
[0106] Step 5, performing downlink power adjustment according to the power adjustment value to determine the downlink power of each preset bandwidth, for example, the downlink power can be allocated to each RB resource.
[0107] The above technical solution can accurately predict the channel quality of the first time unit of the future power to be adjusted, finely adjust the downlink power, greatly improve the flexibility of downlink power allocation, and improve the service experience of users. At the same time, the intelligent single board is added at the edge side, which can reduce the processing delay while enhancing the edge computing power, and more accurate downlink power control can be achieved.
[0108] Figure 8 A schematic diagram of a model training device provided by an embodiment of the present specification is shown in Figure 8 The model training device 100 includes: The acquisition module 110 is configured to acquire the first channel quality reported by the terminal.
[0109] The prediction module 120 is configured to process the first channel quality to obtain the second channel quality corresponding to at least one first time unit.
[0110] The adjustment module 130 is configured to adjust the second downlink power of the target time unit according to the second channel quality corresponding to the first time unit and the first downlink power, to obtain the adjusted third downlink power.
[0111] The length of the first time unit is less than the minimum reporting period of the first channel quality, and the target time unit is at least one first time unit of the power to be adjusted.
[0112] Optionally, the adjusting module 130 is further configured to determine the power adjustment value of the target time unit according to the second channel quality corresponding to the first time unit and the first downlink power. According to the power adjustment value, the second downlink power of the target time unit is adjusted to obtain the adjusted third downlink power.
[0113] Optionally, the adjusting module 130 is further configured to determine the power adjustment value of the target time unit according to the second channel quality corresponding to the first time unit and the first downlink power by using a reinforcement learning algorithm.
[0114] The power adjustment value is a power adjustment value corresponding to at least one preset bandwidth respectively; the reinforcement learning algorithm is a Q-learning algorithm; a state space of the Q-learning algorithm is a two-dimensional state space corresponding to the at least one preset bandwidth respectively, the two-dimensional state space dimension includes the channel quality reported by the terminal and the corresponding downlink power, an action space of the Q-learning algorithm is a power adjustment value corresponding to the at least one preset bandwidth respectively, and a reward function of the Q-learning algorithm is a weighted average value of the channel quality reported by the terminal corresponding to the at least one preset bandwidth respectively. The preset bandwidth includes any one of the following: A first preset number of resource blocks (RBs); A second preset number of subbands.
[0115] Optionally, the reinforcement learning algorithm further includes a weight of the reward function, and the weight of the reward function includes at least one of the following: A user weight corresponding to the terminal; A service weight corresponding to different service types.
[0116] Optionally, the predicting module 120 is further configured to infer the first channel quality and the corresponding fourth downlink power by using a pre-trained neural network model to obtain the second channel quality corresponding to at least one first time unit output by the neural network model. The loss function of the neural network model includes a first loss function, and the first loss function represents the correlation between the predicted second channel quality and the corresponding first downlink power.
[0117] Optionally, the loss function of the neural network model further includes at least one of a second loss function and a third loss function, the second loss function represents a mean square error between the second channel quality and the corresponding first channel quality, and the third loss function represents a distribution error between the second channel quality and the corresponding first channel quality.
[0118] The device 100 provided by the embodiment of the present application can execute each method in the foregoing method embodiments, and realize the functions and advantages of each method in the foregoing method embodiments, which will not be repeated here.
[0119] Figure 9 The schematic diagram of another downlink power adjustment device provided by an embodiment of the present application is shown in FIG. 1, which comprises a training module 140. Figure 9 The training module 140 is configured to: Step 1, determining the initial value of the Q value table as a preset value, the Q value table being a two-dimensional table composed of a state space and a power adjustment value space; Step 2, determining a first adjustment value from the power adjustment value space according to a first state corresponding to a current first time unit and a preset selection strategy, updating the Q value table corresponding to the first state and the first adjustment value, the first state comprising a channel quality corresponding to the current first time unit and a downlink power; Step 3, taking the next first time unit as the current first time unit, repeating step 2 until the Q value table converges, and obtaining the trained Q value table.
[0120] Optionally, the neural network model is a long short-term memory neural network, and the training module 140 is further configured to: Step 1, obtaining a third channel quality reported by the training terminal at a plurality of third time units and a fifth downlink power corresponding to the third time unit; Step 2, using a plurality of third channel qualities and fifth downlink powers in a preset sliding window except the last third channel quality and the fifth downlink power as training data, training the long short-term memory neural network to be trained; Step 3, obtaining a fourth channel quality predicted by the long short-term memory neural network to be trained according to the training data, determining the result of the loss function according to the last third channel quality, the fourth channel quality and the corresponding fifth downlink power in the preset sliding window, and performing back propagation optimization on the long short-term memory neural network to be trained according to the result of the loss function; Step 4, sliding the preset sliding window backward by a third preset number of third time units, repeating steps 2 and 3 until the result of the loss function meets a preset condition, and obtaining the trained long short-term memory neural network.
[0121] The device 100 provided by the embodiment of the present application can execute each method in the foregoing method embodiments, and realize the functions and advantages of each method in the foregoing method embodiments, which will not be repeated here.
[0122] Figure 10 The hardware structure schematic diagram of the electronic device provided by the embodiment of the present application is shown in FIG. 1, which comprises a training module 140. Figure 10As shown, at the hardware level, the electronic device includes at least one processor, and optionally, an internal bus, a network interface, a memory. The memory can include a memory, such as a high-speed Random-Access Memory (RAM), and can also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device can also include other hardware required by the business.
[0123] The processor, the network interface, and the memory can be connected to each other through the internal bus, which can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0124] The memory stores programs. Specifically, the programs can include program codes including at least one computer operation instruction. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0125] The at least one processor reads the corresponding computer program from the non-volatile memory into the memory and then runs, and forms the device for positioning the target user at the logical level. The at least one processor executes the programs stored in the memory, and specifically executes the method disclosed in the first aspect and realizes the functions and beneficial effects of each method described in the foregoing method embodiments, which will not be repeated here.
[0126] The methods disclosed in the embodiments shown in the first aspect of this disclosure can be applied to at least one processor, or implemented by at least one processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the hardware or by instructions in the form of software within at least one processor. The processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0127] The electronic device can also execute the methods described in the preceding method embodiments and achieve the functions and beneficial effects of the methods described in the preceding method embodiments, which will not be repeated here.
[0128] Of course, in addition to software implementation, the electronic device disclosed herein does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0129] This disclosure also proposes a computer-readable storage medium that stores one or more programs, which, when executed by at least one processor, implement the methods disclosed in the embodiments of the first aspect and achieve the functions and beneficial effects of the methods described in the foregoing method embodiments, which will not be repeated here.
[0130] The computer readable storage medium includes a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.
[0131] Further, the embodiment of the present disclosure also provides a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, when the program instructions are executed by a computer, the following processes are implemented: the method disclosed in the embodiment of the first aspect and the functions and beneficial effects of the methods described in the foregoing method embodiments, which are not described here again.
[0132] In summary, the above only describes the preferred embodiments of the present disclosure, and does not limit the protection scope of the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
[0133] The system, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions. A typical implementation device is a computer. Specifically, the computer may, for example, be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0134] The computer readable medium includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can store information accessible by a computing device. According to the definition herein, the computer readable medium does not include transitory computer readable media, such as modulated data signals and carriers.
[0135] It is also to be noted that the terms "comprising", "including", and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.
[0136] The various embodiments described in this specification are presented as examples. Each example is provided by way of explanation of the overall subject matter and is not a limitation on the overall subject matter. Changes in, or replacements to, parts of certain examples are covered by this specification. The various embodiments described in this specification are presented as examples. Each example is provided by way of explanation of the overall subject matter and is not a limitation on the overall subject matter. Changes in, or replacements to, parts of certain examples are covered by this specification.
Claims
1. A method for downlink power adjustment, the method comprising: The method comprises: obtaining a first channel quality reported by a terminal; processing the first channel quality to obtain a second channel quality corresponding to at least one first time unit; adjusting a second downlink power of a target time unit according to the second channel quality corresponding to the first time unit and a first downlink power to obtain a third downlink power after adjustment; wherein the length of the first time unit is less than the minimum reporting period of the first channel quality, and the target time unit is at least one first time unit whose power is to be adjusted.
2. The method of claim 1, wherein, The method comprises: determining a power adjustment value of the target time unit according to the second channel quality corresponding to the first time unit and the first downlink power; adjusting the second downlink power of the target time unit according to the power adjustment value to obtain the third downlink power after adjustment.
3. The method of claim 2, wherein, The method comprises: adopting a reinforcement learning algorithm to determine the power adjustment value of the target time unit according to the second channel quality corresponding to the first time unit and the first downlink power; wherein the power adjustment value is a power adjustment value corresponding to at least one preset bandwidth respectively; the reinforcement learning algorithm is a Q-learning algorithm; the state space of the Q-learning algorithm is a two-dimensional state space corresponding to at least one preset bandwidth respectively, the dimensions of the two-dimensional state space include the channel quality reported by the terminal and the corresponding downlink power, the action space of the Q-learning algorithm is a power adjustment value corresponding to at least one preset bandwidth respectively, and the reward function of the Q-learning algorithm is a weighted average value of the channel quality reported by the terminal corresponding to at least one preset bandwidth respectively; The preset bandwidth comprises any one of the following: a first preset number of resource blocks (RBs); a second preset number of sub-bands.
4. The method of claim 3, wherein, The reinforcement learning algorithm further comprises a weight of the reward function, and the weight of the reward function comprises at least one of the following: a user weight corresponding to the terminal; a service weight corresponding to different service types.
5. The method of claim 3, wherein, The Q value table corresponding to the Q-learning algorithm is obtained by the following method: Step 1: determining that the initial value of the Q value table is a preset value, and the Q value table is a two-dimensional table composed of the state space and the power adjustment value space; Step 2: determining a first adjustment value from the power adjustment value space according to a first state corresponding to a current first time unit according to a preset selection strategy, updating the Q value table corresponding to the first state and the first adjustment value, and the first state comprises the channel quality and the downlink power corresponding to the current first time unit; Step 3: taking the next first time unit as the current first time unit, repeating the step 2 until the Q value table converges, and obtaining the trained Q value table.
6. The method according to any one of claims 1 to 5, characterized in that, The step of processing the first channel quality to obtain the second channel quality corresponding to at least one first time unit includes: By reasoning about the first channel quality and the corresponding fourth downlink power using a pre-trained neural network model, the second channel quality corresponding to at least one first time unit output by the neural network model is obtained. The loss function of the neural network model includes a first loss function, which characterizes the correlation between the predicted second channel quality and the corresponding first downlink power.
7. The method of claim 6, wherein, The loss function of the neural network model further includes at least one of a second loss function and a third loss function, wherein the second loss function characterizes the mean square error between the second channel quality and the corresponding first channel quality, and the third loss function characterizes the distribution error between the second channel quality and the corresponding first channel quality.
8. The method of claim 6, wherein, The neural network model is a long short-term memory neural network, and the neural network model is trained using the following methods: Step 1: Obtain the third channel quality and the fifth downlink power corresponding to the third time unit reported by the training terminal in multiple third time units; Step 2: Use multiple third channel quality and fifth downlink power data other than the last third channel quality and fifth downlink power within a preset sliding window as training data to train the long short-term memory neural network to be trained; Step 3: Obtain the fourth channel quality predicted by the long short-term memory neural network to be trained based on the training data; determine the result of the loss function based on the last third channel quality of the preset sliding window, the fourth channel quality, and the corresponding fifth downlink power; and perform backpropagation optimization on the long short-term memory neural network to be trained based on the result of the loss function. Step 4: Slide the preset sliding window backward by a third preset number of third time units, and repeat Step 2 and Step 3 until the result of the loss function meets the preset conditions, thereby obtaining the trained long short-term memory neural network.
9. A downlink power adjustment apparatus, characterized by comprising: The device includes: The acquisition module is used to acquire the first channel quality reported by the terminal; The prediction module is used to process the first channel quality to obtain the second channel quality corresponding to at least one first time unit; The determining module is used to adjust the second downlink power of the target time unit according to the second channel quality and the first downlink power corresponding to the first time unit, so as to obtain the adjusted third downlink power; Wherein, the length of the first time unit is less than the minimum reporting period of the first channel quality, and the target time unit is at least one first time unit of the power to be adjusted.
10. An electronic device, comprising: The electronic device includes a memory and a processor. The memory stores computer-executable instructions, which, when executed on the processor, are capable of implementing the steps of the downlink power adjustment method according to any one of claims 1 to 8.
11. A computer-readable storage medium having stored computer-executable instructions therein, wherein, When executed by a processor, the computer-executable instructions are capable of implementing the steps of the downlink power adjustment method according to any one of claims 1 to 8.
12. A computer program product comprising a processing program, characterized in that The processing program is executed by the processor to implement the steps of the downlink power adjustment method of any one of claims 1 to 8.