Wireless Resource Scheduling Method, Training Method and Device for Active Resource Scheduling Model
By using the active resource scheduling model of user-side and base station-side data combined with reinforcement learning algorithms, the problem of insufficient coordination of prediction errors and prediction information in the prior art is solved, and the robustness and efficient utilization of wireless resource scheduling are achieved.
Patent Information
- Application Number
- CN202011117060.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-10-19
AI Technical Summary
The existing wireless resource scheduling methods rely on accurate prediction information and do not consider the impact of prediction error on resource scheduling schemes. There is a lack of a unified coordination mechanism between multiple prediction information, resulting in unstable scheduling performance and waste of resources, and increasing service delay.
User-side data and base-station data are used to predict the performance and user behavior of base stations. Combined with the active resource scheduling model of reinforcement learning algorithm, action parameters are obtained through state parameter analysis to achieve robustness of resource scheduling and ensure that resources can still be utilized efficiently in the presence of prediction errors.
The robustness of resource scheduling in the presence of prediction errors is achieved, resource utilization efficiency is improved, resource waste and service delay are reduced, and the stability of the scheduling method is ensured.
Smart Images

Figure CN114390710B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless technologies, and in particular, to a wireless resource scheduling method, a training method for an active resource scheduling model, and an apparatus. Background Art
[0002] The use of big data and artificial intelligence technologies in wireless networks will greatly improve the performance of the network. The rich data in wireless networks makes it possible to learn "knowledge" from the data and then achieve network optimization. These "knowledge" mainly play a role through the prediction of network performance or user behavior, etc. According to statistics, 90% of network and user behaviors are predictable, such as channel changes and user mobility. Utilizing this predictive information to plan resource allocation in advance, that is, active resource scheduling, will greatly improve the flexibility of resource scheduling, adapt to user behavior on a large time scale, improve resource utilization, and reduce service latency.
[0003] The existing technical solutions for improving the performance of wireless networks using predictive information mainly include:
[0004] 1) Designing a resource scheduling method based on accurate predictive information. This method assumes that accurate predictive information about user mobility, channel state, and user traffic demand is known, and studies the resource scheduling problem within the prediction window, that is, designing a resource scheduling method on the premise that the relevant predictive information meets the accuracy requirements;
[0005] 2) Estimating the average performance thresholds of the network and the base station based on predictive information to guide user scheduling and data transmission.
[0006] However, the existing active resource scheduling methods have the following problems:
[0007] 1) Relying on accurate predictive information, without considering the impact of prediction errors on the resource scheduling scheme, and it is difficult to obtain accurate predictive information in actual systems. Using predictive information with errors will cause the instability of the performance of the scheduling method and cannot fully exert the capabilities of the scheduling method;
[0008] 2) There is a lack of a unified coordination mechanism among multiple predictive information. The inaccuracy of a certain predictive information will affect the overall scheduling performance, resulting in resource waste and increasing service latency;
[0009] 3) When using AI algorithms to implement resource scheduling, there is a lack of a dynamic scheduling algorithm selection mechanism based on scenarios and cost-benefit evaluation. In some scenarios where the performance gain of AI algorithms compared to traditional algorithms is not obvious, choosing an implementation method of AI algorithms with a large computational requirement is likely to cause resource waste. Summary of the Invention
[0010] The object of the technical solution of the present invention is to provide a wireless resource scheduling method, a training method and device for an active resource scheduling model, so as to solve the problems in the prior art that active resource scheduling depends on accurate prediction information, but does not consider the impact of prediction errors on the resource scheduling scheme, and there is a lack of a unified coordination mechanism among multiple prediction information, inaccurate prediction information will affect the overall scheduling performance, resulting in unstable performance of the scheduling method, as well as resource waste and increased service delay.
[0011] An embodiment of the present invention provides a wireless resource scheduling method, which includes:
[0012] Obtain prediction result information according to user-side data and base station-side data;
[0013] Obtain the state parameters of the active resource scheduling model according to the cumulative data transmission volume of the user and the prediction result information;
[0014] Analyze the state parameters through an active resource scheduling model adopting a reinforcement learning algorithm to obtain the action parameters of resource scheduling;
[0015] Output a resource scheduling result according to the action parameters.
[0016] Optionally, in the wireless resource scheduling method, the prediction result information represents the prediction result of the target parameter in each resource configuration time unit in the time window in a time series.
[0017] Optionally, in the wireless resource scheduling method, the prediction result information includes at least one of the following prediction results of target parameters:
[0018] User mobility, achievable spectral efficiency of the user accessing the base station, and available resources of the base station.
[0019] Optionally, in the wireless resource scheduling method, the action parameters represent the users accessed and the amount of resources configured in each resource configuration time unit in the time window in a time series.
[0020] Optionally, in the wireless resource scheduling method, the outputting a resource scheduling result according to the action parameters includes:
[0021] When it is determined according to the action parameters that the amount of resources configured in each resource configuration time unit is less than a preset threshold, passive resource scheduling is performed;
[0022] When it is determined according to the action parameters that the amount of resources configured in each resource configuration time unit is greater than or equal to the preset threshold, active resource scheduling is performed.
[0023] Optionally, in the wireless resource scheduling method, the state parameters include at least one of the following information:
[0024] The predicted value of user access to the base station, the available bandwidth of the base station, the achievable spectral efficiency of user access to the base station, and the cumulative user transmission data volume.
[0025] An embodiment of the present invention further provides a training method for an active resource scheduling model, which includes:
[0026] Obtain prediction result information according to user-side data and base station-side data;
[0027] Obtain the state parameters of the active resource scheduling model according to the cumulative user data transmission volume and the prediction result information;
[0028] Perform model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to satisfy, and obtain an active resource scheduling model that satisfies the objective function.
[0029] Optionally, in the training method, the performing model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to satisfy, and obtaining an active resource scheduling model includes:
[0030] Through the Actor network, according to the current state parameter s t and the constraint conditions, calculate the current personal reward parameter r(s t , a t ), and determine the current action parameter a according to the current Q value t ; where the Q value is the evaluation value of the Actor network policy;
[0031] Through the Critic network, according to the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 , calculate the updated Q value;
[0032] In the case where it is determined that the objective function is not satisfied according to the updated Q value, the current action parameter a t and the data transmission performance parameters, through the Actor network, according to the state parameter s at the next moment t+1 calculate the personal reward parameter r(s at the next moment t+1 , a t+1 ), and determine the action parameter a at the next moment according to the updated Q value t+1 ;
[0033] In the case where it is determined that the objective function is not satisfied according to the updated Q value, the current action parameter a tWhen it is determined that the data transmission performance parameter satisfies the objective function, the model training is ended.
[0034] Optionally, in the training method, the method further includes:
[0035] Storing the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 into the experience pool;
[0036] Among them, the Critic network of the critic obtains the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 .
[0037] Optionally, in the training method, the prediction result information represents the prediction results of the target parameters in each resource configuration time unit in the time window in a time series.
[0038] Optionally, in the training method, the prediction result information includes at least one of the following prediction results of the target parameters:
[0039] User mobility, achievable spectral efficiency of the user accessing the base station, and available resources of the base station.
[0040] Optionally, in the training method, the action parameter represents the users accessed and the amount of resources configured in each resource configuration time unit in a time series.
[0041] Optionally, in the training method, the state parameter includes at least one of the following information:
[0042] Predicted value of user access to the base station, available bandwidth of the base station, achievable spectral efficiency of the user accessing the base station, and cumulative amount of data transmitted by the user.
[0043] Optionally, in the training method, the constraint conditions include at least one of the following:
[0044] The resources allocated by the base station do not exceed the total amount of data transmission;
[0045] The resources allocated to each user in each resource configuration time unit meet the user's requirements;
[0046] The resources allocated to the user meet the user's quality of service QoS requirements.
[0047] An embodiment of the present invention further provides a radio resource scheduling device, which includes:
[0048] A first prediction acquisition module, configured to obtain prediction result information according to user-side data and base station-side data;
[0049] A first state conversion module, configured to obtain state parameters of an active resource scheduling model according to the user's cumulative data transmission volume and the prediction result information;
[0050] A resource scheduling module, configured to analyze the state parameters through an active resource scheduling model adopting a reinforcement learning algorithm to obtain action parameters of resource scheduling;
[0051] An output module, configured to output a resource scheduling result according to the action parameters.
[0052] An embodiment of the present invention further provides a training device for an active resource scheduling model, which includes:
[0053] A second prediction acquisition module, configured to obtain prediction result information according to user-side data and base station-side data;
[0054] A second state conversion module, configured to obtain state parameters of an active resource scheduling model according to the user's cumulative data transmission volume and the prediction result information;
[0055] A training module, configured to perform model training according to data transmission performance parameters, constraint conditions that an active resource scheduling model needs to meet, and the state parameters to obtain an active resource scheduling model that meets the objective function.
[0056] An embodiment of the present invention further provides a processing device, which includes: a processor, a memory, and a program stored on the memory and executable on the processor, and when the program is executed by the processor, it implements the radio resource scheduling method described in any one of the above.
[0057] An embodiment of the present invention further provides a processing device, which includes: a processor, a memory, and a program stored on the memory and executable on the processor, and when the program is executed by the processor, it implements the training method of the active resource scheduling model described in any one of the above.
[0058] An embodiment of the present invention further provides a control system, which includes the processing device described above and the processing device described above.
[0059] An embodiment of the present invention also provides a readable storage medium, wherein a program is stored on the readable storage medium, and when the program is executed by a processor, the steps in the wireless resource scheduling method described in any one of the above are implemented, or the steps in the training method of the active resource scheduling model described in any one of the above are implemented.
[0060] At least one of the above technical solutions of the present invention has the following beneficial effects:
[0061] By using the wireless resource scheduling method described in the embodiment of the present invention, relevant predictions of base station performance and user behavior are made by using user-side data and base station-side data to achieve active resource scheduling, and the robustness of resource scheduling is achieved by using an active resource scheduling model adopting a reinforcement learning algorithm. In this way, even in the case of prediction errors, the data required for resource scheduling can still be determined, realizing the efficient utilization of resources, and solving the problems in the prior art that rely on accurate prediction information for active resource scheduling, but do not consider the impact of prediction errors on the resource scheduling scheme, and there is a lack of a unified coordination mechanism among multiple prediction information, and inaccurate prediction information will affect the overall scheduling performance, resulting in unstable performance of the scheduling method, as well as resource waste and increased service delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a schematic flowchart of the wireless resource scheduling method described in the embodiment of the present invention;
[0063] Figure 2 It is a schematic diagram of the calculation principle of the basic model adopting the reinforcement learning algorithm;
[0064] Figure 3 It is a process architecture diagram of the wireless resource scheduling method described in the embodiment of the present invention;
[0065] Figure 4 It is a timing diagram of the representation of prediction result information;
[0066] Figure 5 It is a timing diagram of the representation of action parameters;
[0067] Figure 6 It is a schematic diagram of the resource scheduling process;
[0068] Figure 7 It is a schematic flowchart of the training method of the active resource scheduling model described in the embodiment of the present invention;
[0069] Figure 8 It is a process architecture diagram of the training method of the active resource scheduling model described in the embodiment of the present invention;
[0070] Figure 9 It is a specific process schematic diagram of the training method of the active resource scheduling model in the embodiment of the present invention;
[0071] Figure 10 Schematic diagram of the structure of a radio resource scheduling device according to one embodiment of the present invention;
[0072] Figure 11 Schematic diagram of the structure of a radio resource scheduling device according to another embodiment of the present invention;
[0073] Figure 12 Schematic diagram of the structure of a processing device according to one embodiment of the present invention;
[0074] Figure 13 Schematic diagram of the structure of a processing device according to another embodiment of the present invention;
[0075] Figure 14 Schematic diagram of the structure of the control system according to the embodiment of the present invention. Detailed implementation manners
[0076] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0077] To solve the problems in the prior art that active resource scheduling depends on accurate prediction information, but does not consider the impact of prediction errors on the resource scheduling scheme, and there is a lack of a unified coordination mechanism between multiple prediction information, inaccurate prediction information will affect the overall scheduling performance, resulting in unstable performance of the scheduling method, as well as resource waste and increased service delay, the embodiments of the present invention provide a radio resource scheduling method, which uses user-side data and base station-side data to perform relevant predictions on base station performance and user behavior, realizes active resource scheduling, and uses an active resource scheduling model based on a reinforcement learning algorithm for analysis to ensure the robustness of prediction errors, so that the radio resource scheduling has active adaptability to environmental changes and can formulate accurate scheduling strategies according to changes in user behavior and base station performance.
[0078] As Figure 1 shown, one embodiment of the present invention provides a radio resource scheduling method, including:
[0079] S110, obtaining prediction result information according to user-side data and base station-side data;
[0080] S120, obtaining state parameters of an active resource scheduling model according to the cumulative data transmission volume of the user and the prediction result information;
[0081] S130, analyzing the state parameters through an active resource scheduling model using a reinforcement learning algorithm to obtain action parameters for resource scheduling;
[0082] S140, outputting a resource scheduling result according to the action parameters.
[0083] By using the wireless resource scheduling method described in this embodiment, relevant predictions of base station performance and user behavior are made using user-side data and base station-side data to achieve proactive resource scheduling, and the proactive resource scheduling model using the reinforcement learning algorithm is used to achieve the robustness of resource scheduling. In this way, even in the presence of prediction errors, the data required for resource scheduling can still be determined, realizing the efficient utilization of resources, and solving the problems in the prior art that rely on accurate prediction information for proactive resource scheduling, but do not consider the impact of prediction errors on the resource scheduling scheme, and there is a lack of a unified coordination mechanism among multiple prediction information. The inaccuracy of prediction information will affect the overall scheduling performance, resulting in unstable performance of the scheduling method, as well as resource waste and increased service delay.
[0084] As Figure 2 shown is a schematic diagram of the calculation principle of the basic model using the reinforcement learning algorithm.
[0085] The basic model using the reinforcement learning algorithm consists of three basic elements: State is a description of the environment; Action is a description of the agent's behavior; Reward is a mapping from a state (or state-action pair) to a reinforcement signal, that is, an evaluation of the previously selected action.
[0086] The agent obtains the current state s t and the reward value r t from the environment, and outputs the selected action a t through the internal state value update mechanism. Under the action of the action a t , the environment switches to a new state s t+1 , and at the same time generates a reinforcement signal r t+1 (reward or punishment), which is the immediate reward. The immediate reward is fed back to the agent, and the agent then selects the next action according to the current state of the environment under the principle of ensuring an increasing probability of positive reward for the agent. The selected action a t+1 not only affects the current immediate reward, but also affects the subsequent and final cumulative rewards.
[0087] Based on the above, the key to using the reinforcement learning algorithm is to design the three basic elements of State, Action, and Reward according to the description of the specific problem, and use the above-mentioned cyclic update of the agent and the environment for these elements to finally obtain a learning result that meets the preset conditions.
[0088] Using the above principle, in the wireless resource scheduling method according to an embodiment of the present invention, based on the user-side data and the base station-side data, the obtained prediction result information is used for state value conversion to obtain the state parameters of the proactive resource scheduling model, that is, the design state of the proactive resource scheduling model. Through the model calculation principle of the above reinforcement learning algorithm, the proactive resource scheduling model generates the action parameters for resource scheduling, that is, obtains the resource scheduling result, so as to use the action parameters for resource scheduling.
[0089] Figure 3 It is a process architecture diagram for adopting the wireless resource scheduling method according to an embodiment of the present invention. Combining Figure 1 and Figure 3 as shown, in the wireless resource scheduling method according to an embodiment of the present invention:
[0090] Step S110, obtain prediction result information according to the user-side data and the base station-side data; wherein the prediction result information includes at least one of a user mobility prediction result, an achievable spectral efficiency prediction result of a user accessing a base station, and a base station available resource prediction result;
[0091] Optionally, the user-side data may include the user access base station address, the user access time, and the user residence time, which are used for predicting user mobility to obtain the user mobility prediction result;
[0092] Optionally, the user-side data and / or the base station-side data may include the distance between the user and the base station, which is used for predicting the achievable spectral efficiency of the user accessing the base station to obtain the achievable spectral efficiency prediction result of the user accessing the base station;
[0093] Optionally, the base station-side data may include the resource reservation information for each time slot, which is used for predicting the available resources of the base station to obtain the base station available resource prediction result.
[0094] Step S120, obtain the state parameters of the proactive resource scheduling model according to the user cumulative data transmission volume and the prediction result information; that is, the prediction result information and the user cumulative data transmission volume are used as input data to output the state parameters of the proactive resource scheduling model.
[0095] Step S130, based on the input state parameters, through the inference analysis of the proactive resource scheduling model, output the action parameters for resource scheduling, that is, obtain the resource scheduling parameters.
[0096] In an embodiment of the present invention, optionally, the state parameters include at least one of the following information:
[0097] The user access base station prediction value, the base station available bandwidth, the achievable spectral efficiency of the user accessing the base station, and the user cumulative transmission data volume.
[0098] In an embodiment of the present invention, optionally, to facilitate converting the prediction result information into state parameters of the active scheduling model in step S120, the prediction result information represents the prediction result of the target parameter in each resource allocation time unit of each time window in a time series. Optionally, the measurement result information includes at least one of the following prediction results of the target parameter:
[0099] User mobility, achievable spectral efficiency of the user accessing the base station, and available resources of the base station.
[0100] For example: The prediction result of user mobility is expressed as: φ u ={φ t,u};
[0101] The prediction result of the achievable spectral efficiency of the user accessing the base station is expressed as:
[0102] The prediction result of the available resources of the base station is expressed as: γ u ={γ t,u};
[0103] Wherein, u represents the user ID; t represents the time window, and each time window may include multiple resource allocation time units. Optionally, the resource allocation time unit may be a time slot or other time length values.
[0104] For example, as shown in Figure 4 , the prediction results of the above target parameters are represented in a time series, with the time window as the minimum prediction time unit. Each time window contains H frames, each frame contains K time slots, and the time slot set of the jth frame The time slot is represented by t, and t = 1, 2,..., HK, and the duration of each time slot is Δt.
[0105] Wherein, the prediction result φ u of user mobility, the prediction result of the achievable spectral efficiency of the user accessing the base station, and the prediction result γ u of the available resources of the base station respectively include multiple data sequences, that is, corresponding to multiple time series, and each data sequence represents the prediction result of the corresponding target parameter in one of the resource allocation time units in the time window.
[0106] In an embodiment of the present invention, optionally, the data sequence included in the prediction result φ u of user mobility is the base station ID accessed by the user.
[0107] Similarly, in an embodiment of the present invention, in step S130, for the output action parameters of resource scheduling, the users accessed and the amount of resources configured in each resource allocation time unit in the time window are represented in a time series.
[0108] Specifically, as Figure 5 shown, the action parameter can represent the proportion of the resources allocated by the base station to user u out of the configurable resources when user u accesses the base station φ t,u at time slot t. The proportion
[0109] is the same as the structure of the input measurement result information. The resource configuration time unit can be a time slot or other time length values; the time window includes one or more resource configuration time units.
[0110] In the embodiment of the present invention, in step S140, according to the action parameter, a resource scheduling result is output, including:
[0111] When it is determined according to the action parameter that the amount of resources configured for each resource configuration time unit is less than a preset threshold, passive resource scheduling is performed;
[0112] When it is determined according to the action parameter that the amount of resources configured for each resource configuration time unit is greater than or equal to the preset threshold, active resource scheduling is performed.
[0113] As Figure 6 shown in the schematic diagram of the resource scheduling process, after determining the action parameter of the resource scheduling, the resource scheduling process starts to be executed. Starting from step S610, it includes:
[0114] S620, determining the amount of resources configured for each resource configuration time unit according to the action parameter;
[0115] S630, determining whether the amount of resources configured for each resource configuration time unit is less than a preset threshold. When it is less than the preset threshold, S640 is executed; when it is greater than or equal to the preset threshold, S650 is executed;
[0116] S640, performing passive resource scheduling;
[0117] S650, performing active resource scheduling;
[0118] S660, scheduling radio resources;
[0119] S670, ending.
[0120] By adopting the wireless resource scheduling method described in the embodiments of the present invention, using user-side data and base station-side data to perform relevant predictions on base station performance and user behavior, realizing proactive resource scheduling, and using an active resource scheduling model based on a reinforcement learning algorithm to achieve the desired results, so as to ensure the robustness of prediction errors, thereby enabling wireless resource scheduling to have proactive adaptability to environmental changes, and being able to formulate precise scheduling strategies according to changes in user behavior and base station performance, thus solving the problems in the prior art that rely on accurate prediction information for proactive resource scheduling, but do not consider the impact of prediction errors on resource scheduling schemes, and there is a lack of a unified coordination mechanism among multiple prediction information, inaccurate prediction information will affect the overall scheduling performance, resulting in unstable performance of the scheduling method, as well as resource waste and increased service delay.
[0121] On the other hand, an embodiment of the present invention further provides a training method for an active resource scheduling model. This training method can be used to train the active resource scheduling model in the above wireless resource scheduling method to obtain a robust active resource scheduling model in the wireless resource scheduling method.
[0122] As Figure 7 shown, the training method for the active resource scheduling model described in the embodiments of the present invention includes:
[0123] S710, obtaining prediction result information according to user-side data and base station-side data;
[0124] S720, obtaining state parameters of the active resource scheduling model according to the cumulative data transmission volume of the user and the prediction result information;
[0125] S730, performing model training according to the data transmission performance parameters, constraint conditions that the active resource scheduling model needs to meet, and the state parameters to obtain an active resource scheduling model that meets the objective function.
[0126] Figure 8 It is a process architecture diagram of the training method for the active resource scheduling model described in the embodiments of the present invention. Combining Figure 7 and Figure 8 shown, in the training method described in the embodiments of the present invention:
[0127] In step S710, obtaining prediction result information according to user-side data and base station-side data; where the prediction result information includes at least one of a user mobility prediction result, an achievable spectral efficiency prediction result of a user accessing a base station, and a base station available resource prediction result;
[0128] Optionally, the user-side data may include the base station address accessed by the user, the user access time, and the user residence time, which are used to perform user mobility prediction to obtain a user mobility prediction result;
[0129] Optionally, the user-side data and / or the base station-side data may include the distance between the user and the base station, which is used for predicting the achievable spectral efficiency of the user accessing the base station, and obtaining the prediction result of the achievable spectral efficiency of the user accessing the base station;
[0130] Optionally, the base station-side data may include the resource reservation information for each time slot, which is used for predicting the available resources of the base station and obtaining the prediction result of the available resources of the base station.
[0131] Step S720: Obtain the state parameters of the active resource scheduling model according to the cumulative data transmission volume of the user and the prediction result information; that is, the prediction result information and the cumulative data transmission volume of the user are used as input data, and the state parameters of the active resource scheduling model are output.
[0132] S730: Perform model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to satisfy, and obtain the active resource scheduling model that satisfies the objective function.
[0133] Adopting this training method, in step S730, the Deep Deterministic Policy Gradient algorithm combines Deep Q-learning (DQN) with the actor-critic reinforcement learning architecture, which can handle the characteristics of continuous action spaces and continuous state spaces, and uses the Deep Deterministic Policy Gradient algorithm to implement the training of the active resource scheduling model.
[0134] Among them, in step S730, as Figure 8 shown, according to the data transmission performance parameters and state parameters that the active resource scheduling model needs to satisfy, use the DDPG network to perform model training.
[0135] Optionally, the data transmission performance parameters are used to configure the wireless transmission requirements that need to be satisfied, and the data transmission performance parameters include, but are not limited to, only the user QoS guarantee index, the user mobility index, and the user upper limit, etc.
[0136] In the training method described in the embodiments of the present invention, the input parameters for model training include prediction result information and data transmission performance parameters.
[0137] The structure of the prediction result information is the same as that of the prediction result information input during the wireless resource scheduling method. The prediction result information represents the prediction result of the target parameter in each resource configuration time unit in each time window in a time series. Optionally, the measurement result information includes at least one of the following prediction results of the target parameter:
[0138] User mobility, the achievable spectral efficiency of the user accessing the base station, and the available resources of the base station.
[0139] For example, the prediction result of user mobility is expressed as: φ u ={φ t,u};
[0140] The prediction result of the achievable spectral efficiency of the user accessing the base station is expressed as:
[0141] The prediction result of the available resources of the base station is expressed as: γ u ={γ t,u};
[0142] Wherein, u represents the user ID; t represents the time window, and each time window can include multiple resource configuration time units. Optionally, the resource configuration time unit can be a time slot or other time length values.
[0143] For example, as shown in Figure 4 The prediction results of the above target parameters are represented in the form of a time series, with the time window as the minimum prediction time unit. Each time window contains H frames, each frame contains K time slots, and the time slot set of the jth frame is The time slot is represented by t, and t = 1, 2,..., HK, and the duration of each time slot is Δt.
[0144] Among them, the prediction result φ of user mobility u , the prediction result of the achievable spectral efficiency of the user accessing the base station and the prediction result γ of the available resources of the base station u respectively include multiple data sequences, that is, corresponding to multiple time series, and each data sequence represents the prediction result of the corresponding target parameter in one of the resource configuration time units in the time window.
[0145] In an embodiment of the present invention, optionally, the data sequence included in the prediction result φ of user mobility u is the base station ID accessed by the user.
[0146] Similarly, in an embodiment of the present invention, in step S720, for the action parameters of the output resource scheduling, the users accessed and the amount of resources configured in each resource configuration time unit in the time window are represented in the form of a time series.
[0147] Specifically, as shown in Figure 5 The action parameter can represent the proportion of the resources allocated by the base station for user u to the configurable resources t,u when user u accesses the base station φ at time slot t
[0148] Having the same structure as the input measurement result information, the resource allocation time unit can be a time slot or other time length values; the time window includes one or more resource allocation time units.
[0149] In the embodiments of the present invention, as Figure 8 shown, the DDPG network consists of a Critic network and an Actor network. The Critic network evaluates the strategy of the Actor network by learning the historical data in the experience pool and updates Q(s, a), where Q(s, a) is also the evaluation value of the Actor network strategy; combined with Figure 2 , the Actor network formulates the behavior strategy of the agent (base station) and gives the action parameters of the agent according to the state parameters. Then the Actor network stores the current state parameter s t , action parameter a t , personal reward parameter r t , and the state parameter s t+1 at the next moment into the experience pool. Among them, the experience pool is a storage database for storing the above-mentioned parameters in each training process during the training process of the active resource scheduling model.
[0150] Using the DDPG network with the above functions, combined with Figure 8 and Figure 9 , the specific process of the training method of the active resource scheduling model includes:
[0151] S910, start;
[0152] S920, collect data transmission performance parameters; optionally, including: user QoS, user number limit, mobility requirements, etc.;
[0153] S930, collect user-side data and base station-side data;
[0154] S940, obtain prediction result information according to the user-side data and the base station-side data; where the prediction result information includes at least one of the user mobility prediction result, the reachable spectral efficiency prediction result of the user accessing the base station, and the available resource prediction result of the base station;
[0155] S950, obtain the state parameter s t of the active resource scheduling model according to the user cumulative data transmission volume and the prediction result information;
[0156] S960, send the state parameter to the Actor network in the DDPG network, calculate the personal reward parameter r(s t , a t ), and give the action parameter a t (resource scheduling result) according to the Q value;
[0157] S970, store the current s t , a t , r(s t , a t ), and the state parameter s t+1 at the next moment into the experience pool;
[0158] S980, the Critic network in the DDPG network updates the Q value according to s t , a t , r(s t , a t ), and the state parameter s t+1 at the next moment;
[0159] S990, determine whether the target function is satisfied according to the data transmission performance parameters, the updated Q value, and the current action parameters required by the active resource scheduling model. If the target function is not satisfied, return to step S930; if the target function is satisfied, execute step S991;
[0160] S991, obtain the robust active resource scheduling model;
[0161] S992, end the model training.
[0162] Therefore, in the embodiment of the present invention, in step S730, the model training is performed according to the data transmission performance parameters required by the active resource scheduling model and the state parameters to obtain an active resource scheduling model that satisfies the target function conditions, including:
[0163] Through the Actor network, according to the current state parameter s t , the constraint conditions calculate the current personal reward parameter r(s t , a t ), and determine the current action parameter a t according to the current Q value;
[0164] Through the Critic network, according to the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ), and the state parameter s t+1 at the next moment, calculate and update the Q value;
[0165] In the case where it is determined that the target function is not satisfied according to the updated Q value, the current action parameter a t and the data transmission performance parameters, through the Actor network, according to the state parameter s t+1Calculate the personal reward parameter r(s t+1 , a t+1 ) at the next moment, and determine the action parameter a t+1 at the next moment according to the updated Q value;
[0166] When it is determined that the objective function is satisfied according to the updated Q value, the current action parameter a t and the data transmission performance parameter, it is determined that the model training is completed. The training method described in the embodiments of the present invention uses the DDPG reinforcement learning algorithm based on the model to make decisions for active resource scheduling. DDPG combines deep Q-learning (DQN) with the actor-critic reinforcement learning architecture and can handle continuous action spaces and continuous state spaces.
[0167] The DDPG algorithm is implemented using the DQN mechanism and needs to design the objective function, constraint conditions, and basic elements of reinforcement learning, including state parameters, action parameters, and personal reward parameters, that are satisfied by the optimization problem according to specific application problems. In the embodiments of the present invention, the above design of the active resource scheduling method is specifically described as follows:
[0168] 1) Objective function
[0169] For ease of calculation, an auxiliary variable T u =(T t,u , T t,u ∈{0, 1}) is introduced. Among them, T t,u =1 means that user u still has data to transmit in time slot t; T t,u =0 means that user u has completed data transmission. The service delay of user u can be expressed as the first-order norm of T u , and the objective function can be expressed as:
[0170]
[0171] 2) Constraint conditions
[0172] According to the configured data transmission performance parameters, including the mobility of users, QoS requirements, and total data transmission volume, etc., the constraint conditions that the active resource scheduling model needs to satisfy are determined as:
[0173] a) Among them indicates the proportion of the resources allocated by the base station for user u when user u accesses the base station φ t,u at time slot t in its configurable resources ; this constraint condition is used to indicate that the allocated resources do not exceed the total amount;
[0174] b) This constraint condition indicates that the resources allocated to each user per time slot meet the user's requirements;
[0175] c) This constraint indicates that the resources allocated to the user meet its QoS requirements.
[0176] Among them, the maximum transmission rate of user u at time slot t is c. t,u ; Among them
[0177] b t,u represents the percentage of the data volume that user u needs to transmit at least in time slot t in its total data demand; represents the available bandwidth of the base station; γ t,u represents the achievable spectral efficiency of the user accessing the base station;
[0178] b′ t,u ∈[0, 1] represents the percentage of the cumulative transmission volume of this user in its total data demand.
[0179] 3) State parameters
[0180] The state parameter S t,u includes four elements. In addition to the predicted value φ of the user accessing the base station t,u , the available bandwidth of the base station , the achievable spectral efficiency γ of the user accessing the base station t,u and other prediction results, it also includes the percentage b′ of the user's cumulative transmitted data volume in its total data demand t,u , which is used to describe the transmission completion situation of the user service.
[0181]
[0182] Among them,
[0183] Among them, b′ t+1,u is the percentage of the user's cumulative transmitted data volume in its total data demand at the next moment.
[0184] 4) Action parameters
[0185] The action parameter a t,u is the percentage of the available resources of the base station allocated to each user in each time slot:
[0186]
[0187] 5) Personal reward parameter
[0188] represents the personal reward r t obtained by user u after taking the action parameter a t,u under the state parameter S u (S t , a t ) and the global reward r(St , a t ), where the agent aims to maximize the long - term global reward. Define the individual reward as:
[0189] r u (S t , a t ) = η u (b′ t+1,u - 1)+p1
[0190] where η u is the user adaptation factor, describing the user's mobility intensity, is the average cell residence time of user u.
[0191] The global reward is:
[0192]
[0193] where p1 is the penalty for violating the above - mentioned constraints.
[0194] According to the above, in the embodiment of the present invention, in the training method, when performing model training, an active resource scheduling model that satisfies the objective function is obtained.
[0195] In the embodiment of the present invention, the constraints for model training include at least one of the following:
[0196] The resources allocated by the base station do not exceed the total data transmission volume;
[0197] For each resource configuration time unit, the resources allocated to each user meet the user's requirements;
[0198] The resources allocated to the user meet the user's quality of service (QoS) requirements.
[0199] In the wireless resource scheduling method and the training method of the active resource scheduling model according to the embodiment of the present invention, the prediction information is used to implement active resource scheduling, and at the same time, the reinforcement learning active scheduling method is adopted to achieve the robustness of the prediction accuracy, ensuring that the efficient utilization of resources can still be achieved in the presence of prediction errors; in addition, the data required for resource scheduling is determined, and the conversion format of the prediction data to the input of the AI model is realized; and the basic elements of reinforcement learning suitable for the robust active resource scheduling problem are designed.
[0200] The embodiment of the present invention also provides a wireless resource scheduling device, as Figure 10 shown, including:
[0201] The first prediction acquisition module 1010 is used to obtain prediction result information according to the user - side data and the base - station - side data;
[0202] The first state conversion module 1020 is configured to obtain state parameters of the active resource scheduling model according to the user's cumulative data transmission volume and the prediction result information;
[0203] The resource scheduling module 1030 is configured to analyze the state parameters through an active resource scheduling model that adopts a reinforcement learning algorithm to obtain action parameters for resource scheduling;
[0204] The output module 1040 is configured to output a resource scheduling result according to the action parameters.
[0205] Optionally, for the wireless resource scheduling device, the prediction result information represents the prediction results of target parameters in each resource configuration time unit in the time window in a time series.
[0206] Optionally, for the wireless resource scheduling device, the prediction result information includes at least one of the following prediction results of target parameters:
[0207] User mobility, achievable spectral efficiency of the user accessing the base station, and available resources of the base station.
[0208] Optionally, for the wireless resource scheduling device, the action parameters represent the users accessed and the amount of resources configured in each resource configuration time unit in the time window in a time series.
[0209] Optionally, for the wireless resource scheduling device, the output module 1040 outputs a resource scheduling result according to the action parameters, including:
[0210] When it is determined according to the action parameters that the amount of resources configured in each resource configuration time unit is less than a preset threshold, passive resource scheduling is performed;
[0211] When it is determined according to the action parameters that the amount of resources configured in each resource configuration time unit is greater than or equal to the preset threshold, active resource scheduling is performed.
[0212] Optionally, for the wireless resource scheduling device, the state parameters include at least one of the following information:
[0213] User access base station prediction value, available bandwidth of the base station, achievable spectral efficiency of the user accessing the base station, and user cumulative transmission data volume.
[0214] An embodiment of the present invention further provides a training device for an active resource scheduling model, as Figure 11 , including:
[0215] The second prediction acquisition module 1110 is configured to obtain prediction result information according to user-side data and base-station-side data;
[0216] The second state conversion module 1120 is configured to obtain state parameters of the active resource scheduling model according to the cumulative data transmission volume of the user and the prediction result information;
[0217] The training module 1130 is configured to perform model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to satisfy, so as to obtain an active resource scheduling model that satisfies the objective function.
[0218] Optionally, in the training device, the training module 1130 performs model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to satisfy, so as to obtain an active resource scheduling model, including:
[0219] Through the Actor network, according to the current state parameter s t and the constraint conditions, calculate the current personal reward parameter r(s t , a t ), and determine the current action parameter a according to the current Q value t ; where the Q value is the evaluation value of the Actor network policy;
[0220] Through the Critic network, according to the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 , calculate the updated Q value;
[0221] In the case that it is determined that the objective function is not satisfied according to the updated Q value, the current action parameter a t and the data transmission performance parameters, through the Actor network, according to the state parameter s at the next moment t+1 calculate the personal reward parameter r(s at the next moment t+1 , a t+1 ), and determine the action parameter a at the next moment according to the updated Q value t+1 ;
[0222] In the case that it is determined that the objective function is satisfied according to the updated Q value, the current action parameter a t and the data transmission performance parameters, determine that the model training is completed.
[0223] Optionally, in the training device, the training module 1130 is further configured to:
[0224] The current state parameter s t , the current action parameter a t, the current personal reward parameter r(s t , a t ) and the state parameter s t+1 at the next moment are stored in the experience pool;
[0225] Among them, the Critic network of the reviewer obtains the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s t+1 at the next moment.
[0226] Optionally, in the training device, the prediction result information represents the prediction results of the target parameters in each resource configuration time unit in the time window in a time series.
[0227] Optionally, in the training device, the prediction result information includes at least one of the following prediction results of the target parameters:
[0228] User mobility, achievable spectral efficiency of the user accessing the base station, and available resources of the base station.
[0229] Optionally, in the training device, the action parameter represents the users accessed and the amount of resources configured in each resource configuration time unit in the time window in a time series.
[0230] Optionally, in the training device, the state parameter includes at least one of the following information:
[0231] Predicted value of user access to the base station, available bandwidth of the base station, achievable spectral efficiency of the user accessing the base station, and cumulative user transmission data volume.
[0232] Optionally, in the training device, the constraint conditions include at least one of the following:
[0233] The resources allocated by the base station do not exceed the total data transmission volume;
[0234] The resources allocated to each user in each resource configuration time unit meet the user's requirements;
[0235] The resources allocated to the user meet the user's quality of service QoS requirements.
[0236] Another aspect of the embodiments of the present invention further provides a processing device. In this embodiment, the processing device is a base station, such as Figure 12As shown, it includes: a processor 1201; and a memory 1203 connected to the processor 1201 through a bus interface 1202. The memory 1203 is used to store the programs and data used by the processor 1201 when performing operations. The processor 1201 calls and executes the programs and data stored in the memory 1203.
[0237] Among them, a transceiver 1204 is connected to the bus interface 1202 and is used to receive and send data under the control of the processor 1201. Specifically, the processor 1201 is used to read the programs in the memory 1203 and execute the following processes:
[0238] Obtain prediction result information based on user-side data and base station-side data;
[0239] Obtain the state parameters of the active resource scheduling model based on the user's cumulative data transmission volume and the prediction result information;
[0240] Analyze the state parameters through an active resource scheduling model using a reinforcement learning algorithm to obtain the action parameters of resource scheduling;
[0241] Output a resource scheduling result according to the action parameters.
[0242] Optionally, for the processing device, where the prediction result information represents the prediction results of the target parameters in each resource configuration time unit in the time window in a time series.
[0243] Optionally, for the processing device, where the prediction result information includes at least one of the following prediction results of the target parameters:
[0244] User mobility, achievable spectral efficiency of the user accessing the base station, and available resources of the base station.
[0245] Optionally, for the processing device, where the action parameters represent the users accessed and the amount of resources configured in each resource configuration time unit in the time window in a time series.
[0246] Optionally, for the processing device, where the processor 1201 outputs a resource scheduling result according to the action parameters, including:
[0247] When it is determined according to the action parameters that the amount of resources configured in each resource configuration time unit is less than a preset threshold, passive resource scheduling is performed;
[0248] When it is determined according to the action parameters that the amount of resources configured in each resource configuration time unit is greater than or equal to the preset threshold, active resource scheduling is performed.
[0249] Optionally, the processing device, wherein the state parameter includes at least one of the following information:
[0250] The predicted value of user access to the base station, the available bandwidth of the base station, the achievable spectral efficiency of user access to the base station, and the cumulative user transmission data volume.
[0251] Wherein, in Figure 12 The bus architecture may include any number of interconnected buses and bridges, specifically, various circuits represented by one or more processors represented by processor 1201 and a memory represented by memory 1203 are linked together. The bus architecture can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface provides an interface. The transceiver 1204 can be a plurality of components, that is, including a transmitter and a receiver, and provides a unit for communicating with various other devices on the transmission medium. The processor 1201 is responsible for managing the bus architecture and general processing, and the memory 1203 can store the data used by the processor 1201 when executing operations.
[0252] Another aspect of the embodiments of the present invention further provides a processing device. In this embodiment, the processing device is a base station, as Figure 13 shown, including: a processor 1301; and a memory 1303 connected to the processor 1301 through a bus interface 1302, where the memory 1303 is used to store the programs and data used by the processor 1301 when executing operations, and the processor 1301 calls and executes the programs and data stored in the memory 1303.
[0253] Wherein, the transceiver 1304 is connected to the bus interface 1302 and is used to receive and send data under the control of the processor 1301. Specifically, the processor 1301 is used to read the program in the memory 1303 and execute the following processes:
[0254] Obtain prediction result information according to the user-side data and the base station-side data;
[0255] Obtain the state parameters of the active resource scheduling model according to the cumulative user data transmission volume and the prediction result information;
[0256] Perform model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to meet, and obtain an active resource scheduling model that meets the objective function.
[0257] Optionally, for the processing device, wherein the processor 1301 performs model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to meet, and obtains the active resource scheduling model, including:
[0258] Through the Actor network, according to the current state parameter s t and the constraint conditions, calculate the current personal reward parameter r(s t , a t ), and determine the current action parameter a according to the current Q value t ; where the Q value is the evaluation value of the Actor network policy;
[0259] Through the Critic network, according to the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 , calculate and update the Q value;
[0260] When it is judged that the objective function is not satisfied according to the updated Q value, the current action parameter a t and the data transmission performance parameter, through the Actor network, according to the state parameter s at the next moment t+1 calculate the personal reward parameter r(s at the next moment t+1 , a t+1 ), and determine the action parameter a at the next moment according to the updated Q value t+1 ;
[0261] When it is judged that the objective function is satisfied according to the updated Q value, the current action parameter a t and the data transmission performance parameter, determine that the model training is completed.
[0262] Optionally, for the processing device, where the processor 1301 is further configured to:
[0263] Store the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 into the experience pool;
[0264] Wherein, the Critic network obtains the current state parameter s from the experience pool t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 .
[0265] Optionally, the processing device, wherein the prediction result information represents the prediction results of the target parameters in each resource configuration time unit in the time window in a time series.
[0266] Optionally, the processing device, wherein the prediction result information includes at least one of the following prediction results of the target parameters:
[0267] User mobility, reachable spectral efficiency of the user accessing the base station, and available resources of the base station.
[0268] Optionally, the processing device, wherein the action parameters represent the users accessed and the amount of resources configured in each resource configuration time unit in the time window in a time series.
[0269] Optionally, the processing device, wherein the state parameters include at least one of the following information:
[0270] Predicted value of user access to the base station, available bandwidth of the base station, reachable spectral efficiency of the user accessing the base station, and cumulative amount of data transmitted by the user.
[0271] Optionally, the processing device, wherein the constraint conditions include at least one of the following:
[0272] The resources allocated to the base station do not exceed the total amount of data transmission;
[0273] The resources allocated to each user in each resource configuration time unit meet the user's requirements;
[0274] The resources allocated to the user meet the user's quality of service (QoS) requirements.
[0275] Wherein, in Figure 13 The bus architecture may include any number of interconnected buses and bridges, specifically, various circuits of one or more processors represented by the processor 1301 and the memory represented by the memory 1303 are linked together. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface provides an interface. The transceiver 1304 may be a plurality of elements, that is, including a transmitter and a receiver, and provides a unit for communicating with various other devices on the transmission medium. The processor 1301 is responsible for managing the bus architecture and general processing, and the memory 1303 can store the data used by the processor 1301 when executing operations.
[0276] Another aspect of the embodiments of the present invention further provides a control system, including the processing devices of the above two structures.
[0277] In an embodiment of the present invention, one of the above processing devices is used for training an active resource scheduling model, and the other processing device is used for implementing active resource scheduling for a base station by using the trained active resource scheduling model. Taking the processing device for training the active resource scheduling model as a wireless intelligent manager and the processing device for implementing active resource scheduling as a wireless intelligent controller as an example, the process of resource scheduling by combining these two processing devices is described as follows, such as Figure 14 shown, the active resource scheduling mainly includes the following processes:
[0278] 1) Robust active resource scheduling model training process
[0279] Step 1: The wireless intelligent manager collects network-level (i.e., base station-level) and user-level data.
[0280] The network-level and user-level data at least includes: the load of different QCI-level services of the base station at the cell level, the load of non-real-time services, the historical record of the access cell ID of the user within a certain time window, the moment when the user enters the cell, the residence duration of the user in the cell, the wireless signal quality, wireless signal strength, signal-to-interference-plus-noise ratio, base station transmit power, etc. of the user's serving cell and neighboring cells for each resource allocation unit within the prediction time window;
[0281] Step 2: Based on the collected data, the wireless intelligent manager trains a user mobility prediction model, a user access base station reachable spectral efficiency prediction model, a base station available resource prediction model, and a robust active resource scheduling model.
[0282] 2) Robust active resource scheduling model deployment
[0283] Step 3: The prediction models and the robust active resource scheduling model obtained in Step 2 are sent down and deployed to the wireless intelligent controller.
[0284] 3) Real-time resource scheduling (robust active resource scheduling model inference) process
[0285] Step 4: The wireless intelligent controller subscribes in real time from the base station information such as the wireless signal quality of the user, the historical access cell ID sequence of the user within a period of time, and the resource load of different types of services of the base station;
[0286] Step 5: The user mobility prediction model predicts in real time the access cell ID of the user within the prediction window; the user access base station reachable spectral efficiency prediction model predicts the signal quality of the user accessing the base station; the base station available resource prediction model predicts in real time the available resources of the base station. The above prediction information is used as the input of the robust active resource scheduling model, and through model inference, the percentage of the available resources allocated by the access base station for the user within the prediction time window is obtained;
[0287] Step 6: The wireless intelligent controller sends the resource scheduling result obtained by inferring the robust active resource scheduling model to the base station;
[0288] Step 7: The base station measures in real time the measurement data such as the signal quality RSRP and SINR of the user, and determines the cell to which the user accesses according to the RSRP maximum criterion;
[0289] Step 8: According to the actually configurable resource amount of the cell to which the user accesses, select a resource scheduling method and determine a resource allocation plan. If the robust active resource scheduling method is selected, determine the resource configuration plan according to the resource scheduling result given by the robust active resource scheduling model; otherwise, determine the resource configuration plan according to the passive resource scheduling method;
[0290] Step 9: The base station performs resource scheduling and data transmission according to the resource configuration plan obtained in Step 8.
[0291] Through the above process, using the user-side data and the base-station-side data, perform relevant predictions on the base-station performance and user behavior to achieve active resource scheduling, and use the active resource scheduling model adopting the reinforcement learning algorithm to achieve the robustness of resource scheduling, so that even in the case of prediction errors, the data required for resource scheduling can still be determined to achieve the efficient utilization of resources.
[0292] In addition, a specific embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, wherein when the program is executed by a processor, it implements the steps in the wireless resource scheduling method described in any one of the above, or implements the steps in the training method of the active resource scheduling model described in any one of the above.
[0293] Specifically, this computer-readable storage medium is applied to the above terminal. When applied to the terminal, the execution steps corresponding to the method for reporting smoke alarm are as detailed above and will not be elaborated here.
[0294] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0295] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing unit, or each unit can be physically separate, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a hardware plus software functional unit.
[0296] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit stored in a storage medium includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute some steps of the transceiver method described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0297] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A wireless resource scheduling method, characterized in that, Including: Obtain prediction result information based on user-side data and base station-side data; the prediction result information includes at least one of the following prediction results of target parameters: user mobility, achievable spectral efficiency of user accessing the base station, and available resources of the base station; Obtain the state parameters of the active resource scheduling model according to the user cumulative data transmission volume and the prediction result information; wherein, the state parameters include at least one of the following information: predicted value of user accessing the base station, available bandwidth of the base station, achievable spectral efficiency of user accessing the base station, and user cumulative transmission data volume; Analyze the state parameters through an active resource scheduling model using a reinforcement learning algorithm to obtain the action parameters of resource scheduling; wherein, the action parameters represent the users accessed and the amount of resources configured for each resource configuration time unit in the time window in a time series; Output a resource scheduling result according to the action parameters; Wherein, the active resource scheduling model is obtained by model training in the following manner according to the data transmission performance parameters, constraint conditions, and state parameters during model training that need to be satisfied; Through the Actor network, according to the current state parameter s during model training t and the constraint conditions, calculate the current personal reward parameter r(s t , a t ), and determine the current action parameter a according to the current Q value t ; where the Q value is the evaluation value of the Actor network policy; Through the Critic network of the reviewer, according to the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 , calculate and update the Q value; When it is determined that the objective function is not satisfied according to the updated Q value, the current action parameter a t and the data transmission performance parameter, through the Actor network, according to the state parameter s t+1 at the next moment, calculate the personal reward parameter r(s t+1 , a t+1 ) at the next moment, and determine the action parameter a t+1 ; When it is determined that the objective function is satisfied based on the updated Q value, the current action parameter a t and the data transmission performance parameter, it is determined that the model training is completed.
2. The wireless resource scheduling method according to claim 1, wherein The prediction result information represents the prediction results of the target parameters for each resource configuration time unit in the time window in a time series.
3. The wireless resource scheduling method according to claim 1, wherein The outputting the resource scheduling result according to the action parameters includes: When it is determined according to the action parameters that the amount of resources configured for each resource configuration time unit is less than a preset threshold, then perform passive resource scheduling; When it is determined according to the action parameters that the amount of resources configured for each resource configuration time unit is greater than or equal to the preset threshold, then perform active resource scheduling.
4. A training method for an active resource scheduling model, characterized in that, Including: Obtain prediction result information based on user-side data and base station-side data; wherein, the prediction result information includes at least one of the following prediction results of target parameters: user mobility, achievable spectral efficiency of user accessing the base station, and available resources of the base station; Obtain the state parameters of the active resource scheduling model according to the user cumulative data transmission volume and the prediction result information; wherein, the state parameters include at least one of the following information: predicted value of user accessing the base station, available bandwidth of the base station, achievable spectral efficiency of user accessing the base station, and user cumulative transmission data volume; Perform model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to satisfy to obtain an active resource scheduling model that satisfies the objective function; Wherein, the performing model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to satisfy to obtain the active resource scheduling model includes: Through the Actor network, according to the current state parameter s t and the constraint conditions, calculate the current personal reward parameter r(s t , a t ), and determine the current action parameter a according to the current Q value t ; where the Q value is the evaluation value of the Actor network policy; Through the Critic network of the reviewer, according to the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 , calculate and update the Q value; When it is determined that the objective function is not satisfied based on the updated Q value, the current action parameter a t and the data transmission performance parameter, through the Actor network, according to the state parameter s t+1 at the next moment, calculate the personal reward parameter r(s t+1 , a t+1 ) at the next moment, and determine the action parameter a t+1 ; When it is determined that the objective function is satisfied based on the updated Q value, the current action parameter a t and the data transmission performance parameter, it is determined that the model training is completed.
5. The training method according to claim 4, characterized in that The method further includes: Deposit the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 into the experience pool; Among them, the critic network of the reviewer obtains the current state parameter s from the experience pool t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 .
6. The training method according to claim 4, characterized in that The prediction result information represents the prediction results of the target parameters for each resource configuration time unit in the time window in a time series.
7. The training method according to claim 4, wherein The action parameter a t Represents, in time series, the users accessing and the amount of resources configured for each resource allocation time unit in the time window.
8. The training method according to claim 4, characterized in that The constraint conditions include at least one of the following: The resources allocated by the base station do not exceed the total data transmission volume; The resources allocated for each user in each resource configuration time unit meet the user requirements; The resources allocated for the user meet the user's quality of service (QoS) requirements.
9. A radio resource scheduling device, characterized in that, Including: The first prediction acquisition module is configured to obtain prediction result information based on user - side data and base - station - side data; the prediction result information includes at least one of the following prediction results of target parameters: user mobility, achievable spectral efficiency of user access to the base station, and available resources of the base station. The first state conversion module is configured to obtain state parameters of the active resource scheduling model based on the cumulative data transmission volume of the user and the prediction result information; wherein, the state parameters include at least one of the following information: predicted value of user access to the base station, available bandwidth of the base station, achievable spectral efficiency of user access to the base station, and cumulative data transmission volume of the user. The resource scheduling module is configured to analyze the state parameters through an active resource scheduling model using a reinforcement learning algorithm to obtain action parameters for resource scheduling; wherein, the action parameters represent the users accessed and the amount of resources configured for each resource configuration time unit in the time window in a time series. The output module is configured to output a resource scheduling result according to the action parameters. Wherein, the active resource scheduling model is obtained by model training in the following manner according to the data transmission performance parameters, constraint conditions, and state parameters that need to be satisfied. Through the Actor network, according to the current state parameter s during model training t and the constraint conditions, calculate the current personal reward parameter r(s t , a t ), and determine the current action parameter a according to the current Q value t ; where the Q value is the evaluation value of the Actor network policy; Through the Critic network of the reviewer, according to the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 , calculate and update the Q value; When it is determined that the objective function is not satisfied based on the updated Q value, the current action parameter a t and the data transmission performance parameter, through the Actor network, according to the state parameter s t+1 at the next moment, calculate the personal reward parameter r(s t+1 , a t+1 ) at the next moment, and determine the action parameter a t+1 at the next moment according to the updated Q value; When it is determined that the objective function is satisfied based on the updated Q value, the current action parameter a t and the data transmission performance parameter, it is determined that the model training is completed.
10. A training device for an active resource scheduling model, characterized in that, Including: The second prediction acquisition module is configured to obtain prediction result information based on user - side data and base - station - side data; wherein, the prediction result information includes at least one of the following prediction results of target parameters: user mobility, achievable spectral efficiency of user access to the base station, and available resources of the base station. The second state conversion module is configured to obtain state parameters of the active resource scheduling model based on the cumulative data transmission volume of the user and the prediction result information; wherein, the state parameters include at least one of the following information: predicted value of user access to the base station, available bandwidth of the base station, achievable spectral efficiency of user access to the base station, and cumulative data transmission volume of the user. The training module is configured to perform model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to satisfy to obtain an active resource scheduling model that meets the objective function. Wherein, the training module performs model training according to the data transmission performance parameters, constraint conditions, and the state parameters that the active resource scheduling model needs to satisfy to obtain an active resource scheduling model, including: Through the Actor network, according to the current state parameter s t and the constraint conditions, calculate the current personal reward parameter r(s t , a t ), and determine the current action parameter a according to the current Q value t ; where the Q value is the evaluation value of the Actor network policy; Through the Critic network of the reviewer, based on the current state parameter s t , the current action parameter a t , the current personal reward parameter r(s t , a t ) and the state parameter s at the next moment t+1 , calculate and update the Q value; When it is determined that the objective function is not satisfied based on the updated Q value, the current action parameter a t and the data transmission performance parameter, through the Actor network, according to the state parameter s t+1 at the next moment, calculate the personal reward parameter r(s t+1 , a t+1 ) at the next moment, and determine the action parameter a t+1 ; When determining that the model training is completed under the condition that the updated Q value, the current action parameter a t and the data transmission performance parameter satisfy the objective function.
11. A processing device, characterized in that, Including: A processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, it implements the radio resource scheduling method according to any one of claims 1 to 3.
12. A processing device, characterized in that, Including: A processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, it implements the training method of the active resource scheduling model according to any one of claims 4 to 8.
13. A control system, characterized in that, Including the processing device according to claim 11 and the processing device according to claim 12.
14. A readable storage medium, characterized in that, A program is stored on the readable storage medium, and when the program is executed by the processor, it implements the steps in the radio resource scheduling method according to any one of claims 1 to 3, or implements the steps in the training method of the active resource scheduling model according to any one of claims 4 to 8.
Citation Information
Patent Citations
Edge computing task allocation method based on deep Monte Carlo tree search
CN110427261A
Wireless resource prediction allocating, acquiring and neural network training method and equipment
CN111491312A