Comprehensive optimization method for base station dormancy control and computing task distribution based on deep reinforcement learning
By optimizing base station sleep and computing task distribution through deep reinforcement learning, the latency and energy consumption problems of base station sleep control and computing task distribution in high-frequency mobile communication networks are solved, thereby reducing base station energy consumption and optimizing user latency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAZHONG UNIV OF SCI & TECH
- Filing Date
- 2024-12-31
- Publication Date
- 2026-05-01
AI Technical Summary
In high-frequency mobile communication networks, the problems of latency and energy consumption optimization in base station sleep control and computing task distribution have not been effectively solved. Especially in the case of multiple latency, existing methods are difficult to balance user experience and energy consumption.
A multi-user edge computing system is constructed using a deep reinforcement learning-based approach. By iteratively optimizing the relationship between users and base stations and the allocation of computing tasks, and by combining the computational cost functions of user latency and base station energy consumption, the sleep state of the base stations is adjusted to achieve comprehensive optimization.
Without significantly impacting user latency performance, it effectively reduces base station energy consumption, achieving a near-global optimal comprehensive optimization result.
Smart Images

Figure CN119854920B_ABST
Abstract
Description
A comprehensive optimization method for base station sleep control and computation task distribution based on deep reinforcement learning. Technical Field
[0001] This invention belongs to the field of mobile communication technology, and more specifically, relates to a comprehensive optimization method for base station sleep control and computing task distribution based on deep reinforcement learning. Background Technology
[0002] In recent years, with the rapid development of mobile communication networks, leveraging mobile edge computing (MEC) to support highly complex and latency-sensitive applications such as autonomous driving and virtual reality has become a trend. By deploying MEC servers at the network edge (such as base stations), mobile users can offload their computing tasks to nearby MEC servers for rapid processing. Simultaneously, to support these data-intensive and latency-sensitive applications, densely deploying a large number of base stations to shorten the distance between base stations and users, thereby improving user performance in terms of transmission rate, latency, and reliability, has become a major development trend in mobile communications. Against this backdrop, implementing reasonable base station sleep strategies is beneficial for reducing base station energy consumption and lowering operator operating costs. However, base station sleep affects initial access, transmission, and computing latency. Therefore, users need to adjust their computing task distribution strategies based on base station sleep conditions to minimize latency.
[0003] Existing research on base station sleep in edge computing systems primarily focuses on low-frequency mobile communication networks, neglecting the characteristics of high-frequency mobile communication networks. In low-frequency bands, channel state changes are relatively gradual, so traditional base station sleep control methods typically make decisions based on historical channel states. However, in high-frequency networks, channel states fluctuate significantly, leading to substantial variations in transmission rate and latency. Especially when users face multiple latency issues (such as initial access latency, data transmission latency, and computation latency), optimizing these latency issues remains a pressing problem. Therefore, in mobile edge computing (MEC) systems, comprehensively optimizing base station sleep control and computation task distribution has both significant theoretical research value and broad application prospects. Summary of the Invention
[0004] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a comprehensive optimization method for base station sleep control and computing task distribution based on deep reinforcement learning, thereby solving the technical problem of joint optimization of base station sleep and computing task distribution in mobile edge computing scenarios.
[0005] To achieve the above objectives, according to one aspect of the present invention, a comprehensive optimization method for base station sleep control and computation task distribution based on deep reinforcement learning is provided, comprising:
[0006] S1. Construct a multi-user edge computing system consisting of a cloud server and multiple edge nodes, where the edge nodes (EN) consist of edge servers and base stations, and determine the initial sleep state of the base stations.
[0007] S2. Based on the current dormant state of the base station, the system first determines the association between the user and the base station, and then, based on the resource status of the user equipment, edge computing server, and cloud computing server, rationally divides the computing tasks and determines the proportion of sub-tasks allocated to each computing node.
[0008] S3. After task allocation, the system calculates the average latency of users and the total energy consumption of base stations. Using these results, it calculates the cost function and generates a corresponding reward value. The system then updates the parameters of the deep reinforcement learning model based on the reward value, further adjusting the base station's sleep state.
[0009] S4. The system repeatedly executes the process of task allocation and base station sleep state adjustment, and gradually optimizes it through multiple rounds of iteration until the scheme converges, finally obtaining a comprehensive optimization result close to the global optimum.
[0010] Preferably, the association between the user and the base station includes:
[0011] The relationship between UE and EN is determined by This indicates that, to ensure each UE is associated with at most one EN, Constraints must be satisfied .
[0012] Preferably, the task division includes:
[0013] For UEs associated with ENj, the generated tasks can be assigned to ENj, a cloud server, or processed by the UE itself. Each task is divided into several sub-tasks, each corresponding to a different data size. The task allocation ratio is used... , and It indicates. Among them, This represents the proportion of tasks handled by UEk itself. This represents the proportion of tasks processed by ENj. This represents the proportion of tasks processed by cloud servers. (Assuming...) Therefore, the division of tasks needs to satisfy the following constraints: ① ;② And when hour, and This indicates that the task is entirely handled by the user terminal.
[0014] Preferably, user mobility modeling includes:
[0015] To capture UE mobility, it is assumed that their location may remain constant or move randomly within each time slot. Therefore, the system needs to re-optimize user association in each time slot. ) and task allocation ( Specifically, the optimization problem of task allocation and user association needs to be solved based on the UE's location information in the current time slot. When the UE's location changes, the system will re-optimize based on the updated location in the next time slot.
[0016] Preferably, the total computing power required to complete the task includes:
[0017] The size of the input data for the task (such as the file size of a video clip) is The unit is bits, and the total number of CPU cycles required to complete the task generated by each UEk is... .
[0018] Preferably, the user computation latency includes:
[0019] The time (in seconds) required to execute the assigned task locally is: ,in, This indicates the number of CPU cycles required to complete the assigned task. Indicates UE Its computing power.
[0020] Assume each EN (Entity Execution Unit) can execute subtasks received from multiple UEs simultaneously. For fairness, the EN's computing power is equally distributed among all UEs associated with that EN in each time slot. Once the input data for a subtask is uploaded to the EN, the EN executes it immediately. This way, each EN incurs no waiting time.
[0021] The time required to execute a subtask of UEk on ENj is determined by the proportion of the subtasks. Input data size and computational complexity Decision, that is ,in This represents the total number of associated users for ENj; The computing power allocated to each ENj is distributed evenly among all associated users. Therefore, the computing power allocated to UEk is... .
[0022] The cloud server provides a fixed amount of computing power to each UEk. If the ratio is The subtasks are unloaded to ENj, and then forwarded by ENj to the cloud server for execution. Then the execution time of the cloud server is... .
[0023] Preferably, the user transmission delay includes:
[0024] Each edge node (EN) is able to measure the uplink signal-to-interference-plus-noise ratio (SINR) of its associated user equipment (UE), denoted as... The specific process is as follows: at the beginning of each time slot, the UE sends a pilot or beacon signal to the EN; the EN uses the received signal to estimate the channel state information (CSI) and calculate the corresponding SINR value.
[0025] set up Given the access channel bandwidth for each EN, the uplink data rate when UEk is associated with ENj is: ,in This is the load of ENj, i.e., the number of UEs associated with it. UEk only... (At a preset threshold value) it can be associated with ENj. The optional EN set of UEk is as follows: like ,but In addition, there is a limit to the number of UEs that each ENj can serve. It is usually determined by the number of physical channels of the EN.
[0026] For the subtasks of UEk, if the allocation ratio is... and Then the time it takes to transmit to EN is .
[0027] The return link between ENj and the cloud server is a wired connection with a bandwidth of [missing information]. Furthermore, the return bandwidth is evenly distributed among all UEs associated with ENj. Therefore, the total time required to upload the subtask of UEk to the cloud server via ENj is... ,in This refers to the backhaul link transmission time.
[0028] Assuming the amount of output data from the task execution is small, the delay in returning the task results from the cloud server or EN is negligible.
[0029] Preferably, the total user latency includes:
[0030] Consider a scenario where a computational task can be divided into multiple independent subtasks, which are then distributed across local devices, edge nodes (ENs), and cloud servers for parallel processing. For example, in a video object recognition task, the video is segmented into several segments, processed separately on local devices, ENs, and cloud servers.
[0031] The task completion time on the local device is .
[0032] The completion time of the task after it is transmitted from UEk to ENj is ,in For the transmission time uploaded to EN, For EN execution time.
[0033] The task is uploaded to EN via UEk, and then forwarded to the cloud server by EN. The total time for returning the result upon completion is: ,in Including upload and backhaul link transmission time, This refers to the computing time of the cloud server.
[0034] Because subtasks are independent and can be executed in parallel, the overall task completion time depends on the last part completed on the local device, EN, and cloud server. This indicates that if UEk selects ENj as the associated node, then the total task completion delay is... The maximum of the following three: .
[0035] Preferably, the base station energy consumption includes:
[0036] Total power consumption of base station It includes two parts: static power consumption and dynamic power consumption, namely... , This represents the static power consumption of the base station when it is turned on, and is related to the fixed configuration of the base station's hardware equipment, cooling system, etc. This part of the power consumption is independent of the base station's operating state and is the basic energy consumption for base station operation. Dynamic power consumption is determined by the number of user equipment (UE) served by the base station and the power consumption of the active radio frequency chain (RF chain). This indicates the power consumption of a single RF chain. This indicates the number of UEs served by base station j, which is equal to the number of active radio links of the base station.
[0037] Preferably, the system cost function includes:
[0038] The system cost function is the weighted sum of the total user latency and the total base station energy consumption accumulated over multiple time points, specifically: ,in, Indicates weight, Let represent the sleep state variable of base station j (0 indicates sleep, 1 indicates no sleep), and constraints exist. This means that the UE can only communicate with base stations that are not in a dormant state. The association is established. In the base station sleep policy based on deep reinforcement learning, the agent's reward is replaced by the negative value of the valence function.
[0039] Optionally, the optimization objective and constraints for minimizing the cumulative average user latency (equivalent to minimizing the total user latency) and total base station energy consumption over multiple time points are as follows:
[0040]
[0041] According to another aspect of the present invention, a comprehensive optimization system for base station sleep control and computation task distribution based on deep reinforcement learning is provided, comprising:
[0042] The network controller collects various status information from other devices in the network, such as user equipment, edge nodes, and cloud servers. Based on this information and preset optimization objectives, it solves for optimal strategies for base station sleep and computation task distribution, and then distributes these strategies to guide network devices to adjust their behavior in response to new instructions. In the sub-problem of base station sleep strategy based on deep reinforcement learning, the network controller acts as an agent to solve for the sleep strategy. In the sub-problem of computation task allocation, the network controller is also the solution unit for user association and computation task distribution.
[0043] User equipment is used to generate computing tasks and communicate with edge nodes or cloud servers.
[0044] Edge nodes, which are composed of edge servers and base stations, are used to receive task requests from user devices and perform task calculations or forward them to cloud servers.
[0045] Cloud servers are used to receive tasks forwarded by edge nodes and complete highly complex computing tasks.
[0046] According to another aspect of the invention, a computer-readable medium is provided that stores a computer program, which, when executed by a processor, performs the steps of the method described above.
[0047] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:
[0048] The proposed integrated optimization method for base station sleep control and computation task distribution based on deep reinforcement learning decomposes the joint optimization problem into two sub-problems: base station sleep and computation tasks based on deep reinforcement learning. The optimal solution is found through continuous iteration of both sub-problems until convergence. The system first constructs a multi-user edge computing system consisting of a cloud server and multiple edge nodes, including edge servers and base stations, and determines the initial sleep state of the base stations. Then, based on the base station's sleep state, the system determines the association between users and base stations and rationally allocates computation tasks according to equipment and computing resources, determining the task allocation ratio for each computing node. After task allocation, the system calculates the average latency of users and the total energy consumption of the base stations. These results are used to calculate a cost function and generate a reward value, which is used to update the parameters of the deep reinforcement learning model, thereby adjusting the sleep state of the base stations. This process involves multiple rounds of iterative optimization and continuous adjustment until the solution converges, ultimately achieving a near-globally optimal integrated optimization result. Based on the above method, this invention effectively reduces base station energy consumption without significantly affecting user latency performance. Attached Figure Description
[0049] Figure 1 is a flowchart of a comprehensive optimization method for base station sleep control and computing task distribution based on deep reinforcement learning in one embodiment of the present invention;
[0050] Figure 2 is a schematic diagram of a comprehensive optimization system for base station sleep control and computing task distribution based on deep reinforcement learning in one embodiment of the present invention;
[0051] Figure 3 is a schematic diagram of the timeline under the optimal subtask partitioning scheme in one embodiment of the present invention. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0053] Example 1
[0054] Figure 1 shows a flowchart of the steps for comprehensive optimization of base station sleep control and computing task distribution based on deep reinforcement learning in one embodiment of the present invention. The steps are described in detail below.
[0055] S1, construct a multi-user edge computing system consisting of a cloud server and multiple edge nodes, where the edge node (EN) consists of an edge server and a base station, and determine the initial sleep state of the base station.
[0056] Specifically, Figure 2 illustrates a schematic diagram of a comprehensive optimization system for base station sleep control and computation task distribution based on deep reinforcement learning in one embodiment. Consider a multi-user mobile edge computing system consisting of a cloud server and multiple mobile edge computing servers. These mobile edge computing servers are placed next to or integrated into the base station of the wireless cellular network. The combination of the base station and the mobile edge servers is considered as an edge node (EN), and each EN communicates with the cloud server via a backhaul connection. The entire system contains J ENs, indexed as follows: These ENs collectively serve K mobile user equipments (UEs), indexed as follows: The base station is initially set to never sleep.
[0057] S2. Based on the current dormant state of the base station, the system first determines the association between the user and the base station, and then, based on the resource status of the user equipment, edge computing server, and cloud computing server, rationally divides the computing tasks and determines the proportion of sub-tasks allocated to each computing node.
[0058] Specifically, to solve the aforementioned optimization problem, this invention employs a hierarchical strategy, decomposing the problem into two levels. The lower level focuses on determining the optimal computational task allocation ratio based on a given user association scheme; the higher level focuses on optimizing the user association scheme. In the lower-level problem, the goal of this invention is to find a computational task allocation method that minimizes the total user latency. Since the final completion of the user device, edge computing server, and cloud server determines the overall latency, and the completion time of each part directly depends on the proportion of tasks allocated to it (the sum of the proportions of the three is 1), minimizing the total latency hinges on ensuring that these three parts can complete their respective tasks as simultaneously as possible. Therefore, this invention achieves the goal of minimizing user latency by finding an optimal task allocation ratio that allows the subtasks of local processing, edge computing, and cloud computing to be completed almost simultaneously. Based on the optimal task allocation scheme derived from the lower level, the higher-level user association problem is transformed into a linear integer programming problem. At this level, this invention uses a dual decomposition-based method to solve it. This method not only ensures that each user is associated with the base station with the most favorable channel conditions but also promotes load balancing among different base stations, thereby improving the efficiency and performance of the entire system. Through this hierarchical optimization strategy, the present invention can effectively solve the comprehensive optimization challenges of base station sleep control and computing task distribution, thereby maximizing system energy efficiency and optimizing user experience.
[0059] Specifically, the present invention is made by This indicates the association relationship between the UE and the EN. To ensure that each UE is associated with at most one EN, the following constraints must be satisfied. For UEs associated with ENj, the generated tasks can be assigned to ENj, a cloud server, or processed by the UE itself. Each task is divided into several sub-tasks, each corresponding to a different data size. The task allocation ratio is determined by... , and It indicates. Among them, This represents the proportion of tasks handled by UEk itself. This represents the proportion of tasks processed by ENj. This represents the proportion of tasks processed by cloud servers. (Assuming...) Therefore, the task division also needs to satisfy two constraints. Constraint one is... This means that only users associated with the base station can offload tasks to edge nodes or cloud servers for computation; constraint two is... This indicates that the sum of the proportions of the three parts into which the task is divided is 1. And when... hour, and This indicates that the task is entirely handled by the user terminal.
[0060] To further explain, in order to capture UE mobility, it is assumed that their location may remain constant or move randomly within each time slot. Therefore, the system needs to re-optimize user association in each time slot. ) and task allocation ( Specifically, the optimization problem of task allocation and user association needs to be solved based on the UE's location information in the current time slot. When the UE's location changes, the system will re-optimize based on the updated location in the next time slot.
[0061] To elaborate further, the size of the input data for the task (such as the file size of a video clip) is The unit is bits, and the total number of CPU cycles required to complete the task generated by each UEk is... The time (in seconds) required to execute the assigned task locally is... ,in, This indicates the number of CPU cycles required to complete the assigned task. This represents the computing power of UEk. It is assumed that each EN can simultaneously execute subtasks received from multiple UEs. For fairness, the computing power of an EN is equally distributed among all UEs associated with that EN in each time slot. Once the input data for a subtask is uploaded to the EN, the EN executes these subtasks immediately. In this way, each EN incurs no waiting time.
[0062] The time required to execute a subtask of UEk on ENj is determined by the proportion of the subtasks. Number of CPU cycles required to complete the total task Decision, that is ,in This represents the total number of associated users for ENj; The computing power allocated to each ENj is distributed evenly among all associated users. Therefore, the computing power allocated to UEk is... .
[0063] The cloud server provides a fixed amount of computing power to each UEk. If the ratio is The subtasks are unloaded to ENj, and then forwarded by ENj to the cloud server for execution. Then the execution time of the cloud server is... .
[0064] To further explain, each edge node (EN) is able to measure the uplink signal-to-interference-plus-noise ratio (SINR) of its associated user equipment (UE), denoted as... The specific process is as follows: at the beginning of each time slot, the UE sends a pilot or beacon signal to the EN; the EN uses the received signal to estimate the channel state information (CSI) and calculate the corresponding SINR value.
[0065] To further explain, suppose Given the access channel bandwidth for each EN, the uplink data rate when UEk is associated with ENj is: ,in This is the load of ENj, i.e., the number of UEs associated with it. UEk only... (Preset threshold value) can be used with EN Association. The optional EN set of UEk is: like ,but In addition, there is a limit to the number of UEs that each ENj can serve. It is usually determined by the number of physical channels of the EN.
[0066] Based on the above, for the subtasks of UEk, if the allocation ratio is... and Then the time it takes to transmit to EN is .
[0067] To further clarify, let's assume the return link between ENj and the cloud server is a wired connection with a bandwidth of [missing information]. Furthermore, the return bandwidth is evenly distributed among all UEs associated with ENj. Therefore, the total time required to upload the subtask of UEk to the cloud server via ENj is... ,in This refers to the backhaul link transmission time.
[0068] To elaborate further, assuming the amount of output data from the task execution is small, the delay in returning the task results from the cloud server or EN is negligible.
[0069] To further illustrate, consider a scenario where the computational task can be divided into multiple independent subtasks, which are then distributed across local devices, edge nodes (ENs), and cloud servers for parallel processing. For example, in a video object recognition task, the video is segmented into several segments, processed separately on local devices, ENs, and cloud servers.
[0070] Based on the above assumptions, the task completion time on the local device is... The completion time of the task after it is transmitted from UEk to ENj is... ,in For the transmission time uploaded to EN, This refers to the execution time of the EN process. The task is uploaded to EN via UEk, then forwarded by EN to the cloud server, and the total time for returning the result upon completion is [time missing]. ,in Including upload and backhaul link transmission time, This refers to the computing time of the cloud server.
[0071] To further explain, because independent subtasks are used, the total latency of a task is equal to the latency of the last set of subtasks completed on the local device, edge node (EN), and cloud server. Since the sum of the proportions of the three parts into which the computational task is divided is 1, reducing the proportion of one part inevitably leads to an increase in the proportion of at least one of the others. Therefore, as shown in Figure 3, optimal task partitioning can be achieved when the subtasks on the local device, edge node, and cloud server complete almost simultaneously. Based on this, the equation can be derived. This indicates that the completion time of each part is equal, where, This indicates the time required for UEk to execute the tasks assigned to it locally. This represents the time required for subtask j of UEk to execute on the edge node. This indicates the time required for subtask j of UEk to execute on the cloud server. This represents the proportion of tasks allocated to each subtask. Based on the previous derivation of latency, the optimal task allocation ratio for UEk can be calculated. , and By combining the task partitioning ratio with a sum of 1, the optimal task partitioning ratio can be determined, which means that the lower-level problems can be solved.
[0072] In the case of user association, the total task completion latency The maximum of the following three: However, the latency is [missing information] when the user is not associated with the service. Therefore, the goal of high-level issues is... ,
[0073] Equivalent to Since the high-level problem is an integer programming problem, this invention relaxes it and uses a dual decomposition method to obtain an approximate optimal solution. After partially relaxing the constraints of the high-level problem, this invention obtains the Lagrangian function. ,in It is the constrained Lagrange multiplier. Higher-level problems can be solved by... This allows us to obtain an approximate optimal solution. The specific solution process is as follows:
[0074] The first step is to set a reasonable initial value. and .
[0075] The second step is for the UE to select the ENj with the best load and signal quality.
[0076] The third step, after receiving selections from each UE, is to determine user associations, calculate gradients, and adjust the settings using gradient methods. This allows the system to gradually approach the optimal state, and finally updates the traffic load. .
[0077] Fourth, after updating the Lagrange multipliers and traffic load, the EN will broadcast the updated information to nearby UEs.
[0078] Fifth, the UE initiates the next round of EN selection process based on the updated information broadcast until a globally approximate optimal solution is reached.
[0079] S3. After task allocation, the system calculates the average latency of users and the total energy consumption of base stations. Using these results, it calculates the cost function and generates a corresponding reward value. The system then updates the parameters of the deep reinforcement learning model based on the reward value, further adjusting the base station's sleep state.
[0080] Specifically, the total power consumption of the base station It includes two parts: static power consumption and dynamic power consumption, namely... , This represents the static power consumption of the base station when it is turned on, and is related to the fixed configuration of the base station's hardware equipment, cooling system, etc. This part of the power consumption is independent of the base station's operating state and is the basic energy consumption for base station operation. Dynamic power consumption is determined by the number of user equipment (UE) served by the base station and the power consumption of the active radio frequency chain (RF chain). This indicates the power consumption of a single RF chain. This indicates the number of UEs served by base station j, which is equal to the number of active radio links of the base station.
[0081] To further explain, the system cost function is the sum of the total user latency and the total base station energy consumption accumulated over multiple time points, i.e. ,in, Indicates weight, Let represent the sleep state variable of base station j (0 indicates sleep, 1 indicates no sleep), and constraints exist. This means that the UE can only communicate with base stations that are not in a dormant state. Establish a connection.
[0082] To further illustrate, the base station sleep control problem is modeled as a Markov Decision Process (MDP), where the network controller acts as an agent, responsible for deciding the sleep control strategy for base stations within its region. Specifically, the system state in the MDP consists of the following information: the current load of edge computing servers and cloud servers, the channel state of users, the transmission rate of users, and the distribution of user-requested tasks. The agent selects an action based on the current system state, i.e., the decision of all base stations to operate or sleep. The reward at each time step is defined as the negative value of the current cost function. The agent's goal is to maximize the sum of the discounted rewards at the current time step and all future time steps through continuous interaction.
[0083] For solving the Multiplication Table (MDP), a deep reinforcement learning-based approach can be employed. This method constructs an agent capable of interacting with the environment, continuously optimizing its decision-making strategy during training to achieve comprehensive optimization of latency and energy consumption. Specifically, the agent first perceives the current system state, including the load of edge computing servers and cloud servers, user channel states, user transmission rates, and task distribution information. Based on the current state, the agent outputs a base station sleep action through a Deep Q-Network (DQN). After executing the sleep action, the environment calculates user latency and base station energy consumption based on the base station's sleep state, user task allocation, and transmission performance, and generates an immediate reward value based on a cost function, which is fed back to the agent. The agent uses the reward signal to evaluate the merits of the action and updates the parameters of its Deep Q-Network, thereby gradually optimizing the decision-making strategy. Through multiple rounds of state updates and action optimization, the agent learns how to make optimal base station sleep decisions under different network states, ultimately achieving comprehensive optimization of latency and energy consumption.
[0084] S4. The system repeatedly executes the process of task allocation and base station sleep state adjustment, and gradually optimizes it through multiple rounds of iteration until the scheme converges, finally obtaining a comprehensive optimization result close to the global optimum.
[0085] Example 2
[0086] This invention also provides a comprehensive optimization system for base station sleep control and computation task distribution based on deep reinforcement learning, comprising:
[0087] The network controller collects various status information from other devices in the network, such as user equipment, edge nodes, and cloud servers. Based on this information and preset optimization objectives, it solves for optimal strategies for base station sleep and computation task distribution, and then distributes these strategies to guide network devices to adjust their behavior in response to new instructions. In the sub-problem of base station sleep strategy based on deep reinforcement learning, the network controller acts as an agent to solve for the sleep strategy. In the sub-problem of computation task allocation, the network controller is also the solution unit for user association and computation task distribution.
[0088] User equipment is used to generate computing tasks and communicate with edge nodes or cloud servers.
[0089] Edge nodes, which are composed of edge servers and base stations, are used to receive task requests from user devices and perform task calculations or forward them to cloud servers.
[0090] Cloud servers are used to receive tasks forwarded by edge nodes and complete highly complex computing tasks.
[0091] Example 3
[0092] The present invention also relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0093] Specifically, the memory may include high-speed random access memory, as well as non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital (SD) cards, flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.
[0094] Example 4
[0095] This invention provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the method described in the above embodiments of this invention.
[0096] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A comprehensive optimization method for base station sleep control and computation task distribution based on deep reinforcement learning, characterized in that, The method includes the following steps: S1, constructing a multi-user edge computing system consisting of a cloud server and multiple edge nodes, wherein the edge nodes (EN) include edge servers and base stations; S2, setting the initial sleep state of the base station, and according to the current sleep state of the base station, allocating computing tasks to user equipment, edge servers, and cloud servers respectively, and determining the proportion of sub-tasks allocated to each computing node; S3, after the task allocation is completed, calculating the average latency of users and the total energy consumption of the base station, and generating an instantaneous reward value based on the cost function, updating the parameters of the deep reinforcement learning model according to the reward value, and further adjusting the sleep state of the base station; in step S3, the base station energy consumption includes: the total power consumption of the base station. It includes two parts: static power consumption and dynamic power consumption, namely... , This represents the static power consumption when the base station is turned on, which is related to the fixed configuration of the base station's hardware equipment, heat dissipation system, etc. Dynamic power consumption is determined by the number of user equipment (UE) served by the base station and the power consumption of the active radio frequency chain (RF chain). Indicates the power consumption of a single RF chain; The number of UEs served by base station j is equal to the number of active radio links of the base station; in step S3, the cost function is the weighted sum of the total user latency and the total base station energy consumption accumulated over multiple time points, specifically as follows: ,in, Indicates weight, This represents the sleep state variable of base station j; 0 indicates sleep, 1 indicates no sleep, and there are constraints. That is, the UE can only associate with base station j that is not in a dormant state; S4, repeatedly execute task allocation and base station dormant state adjustment, and gradually optimize through multiple rounds of iteration until a comprehensive optimization result close to the global optimum is obtained.
2. The comprehensive optimization method for base station sleep control and computation task distribution based on deep reinforcement learning according to claim 1, characterized in that, Step S2 further includes determining the association between the user and the base station, specifically including: the association between the user (UE) and the edge node (EN) is determined by... This indicates that, in order to ensure that each user (UE) is associated with at most one base station (EN), Constraints must be satisfied 。 3. The comprehensive optimization method for base station sleep control and computation task distribution based on deep reinforcement learning according to claim 2, characterized in that, Step S2, which calculates task allocation, includes the following steps: For a user (UE) associated with ENj, the tasks they generate can be allocated to ENj, the cloud server, or processed by the user (UE) themselves; each task is divided into several sub-tasks, each sub-task corresponding to a different data size; the task allocation ratio is used... 、 and Indicates; among which, This represents the proportion of tasks handled by UEk itself. This represents the proportion of tasks processed by ENj. This represents the proportion of tasks processed by cloud servers; assuming... Therefore, the division of tasks needs to satisfy the following constraints: ① ;② ; and when hour, and This indicates that the task is entirely handled by the user terminal.
4. The comprehensive optimization method for base station sleep control and computation task distribution based on deep reinforcement learning according to claim 3, characterized in that, The time required to execute the assigned task locally is ,in, This indicates the number of CPU cycles required to complete the assigned task. Indicates UE The computing power of each EN is assumed to be able to execute subtasks received from multiple UEs simultaneously. For fairness, the computing power of an EN is equally distributed among all UEs associated with that EN in each time slot. Once the input data for a subtask is uploaded to the EN, the EN will immediately execute these subtasks. The time required to execute a subtask of UEk on ENj is determined by the proportion of the subtasks. Number of CPU cycles required to complete the total task Decision, that is ,in This represents the total number of associated users for ENj; The computing power allocated to each ENj is distributed evenly among all associated users; therefore, the computing power allocated to UEk is... The cloud server provides a fixed amount of computing power to each UEk. If the ratio is The subtasks are unloaded to ENj, and then forwarded by ENj to the cloud server for execution. The execution time of the cloud server is... 。 5. The comprehensive optimization method for base station sleep control and computation task distribution based on deep reinforcement learning according to claim 4, characterized in that, User transmission latency includes: the uplink signal-to-interference-to-noise ratio (SINR) of each edge node (EN) that it can measure for its associated user equipment (UE), denoted as... The specific process is as follows: at the beginning of each time slot, the UE sends a pilot or beacon signal to the EN; the EN uses the received signal to perform channel state information (CSI) estimation and calculates the corresponding SINR value; let... Given the access channel bandwidth for each EN, the uplink data rate when UEk is associated with ENj is: ,in It is the load of ENj, i.e., the number of UEs associated with it; UEk only... When a preset threshold value is used, it can be associated with ENj; the optional EN set for UEk is... like ,but In addition, there is a limit to the number of UEs that each ENj can serve. This is typically determined by the number of physical channels in the EN; for subtasks of UEk, if the allocation ratio is... and The task size is Then the time it takes to transmit to EN is The return link between ENj and the cloud server is a wired connection with a bandwidth of [missing information]. Furthermore, the return bandwidth is evenly distributed among all UEs associated with ENj; therefore, the total time required to upload the subtask of UEk to the cloud server via ENj is... ,in This refers to the backhaul link transmission time.
6. The comprehensive optimization method for base station sleep control and computation task distribution based on deep reinforcement learning according to claim 5, characterized in that, Total user latency includes: Considering a divisible computational task where the task can be divided into multiple independent subtasks and distributed for parallel processing on local devices, edge nodes (EN), and cloud servers; in a video object recognition task, the video is segmented into several segments, processed separately on local devices, EN, and cloud servers; the task completion time on the local device is... The completion time of the task after it is transmitted from UEk to ENj is... ,in For the transmission time uploaded to EN, The execution time for EN is [time]. The task is uploaded to EN via UEk, then forwarded to the cloud server by EN, and the total time for returning the result upon completion is [time]. ,in Including upload and backhaul link transmission time, This refers to the computation time on the cloud server; since subtasks are independent and can be executed in parallel, the overall task completion time depends on the last completed portion on the local device, EN, and the cloud server; let This indicates that if UEk selects ENj as the associated node, then the total task completion delay is... The maximum of the following three: 。 7. A computer-readable medium on which a computer program is stored, the computer program, when executed by a processor, implementing the steps of the method as claimed in any one of claims 1-6.