Base station load balancing method and device, storage medium and electronic equipment
By adjusting the base station resource distribution and model parameters based on historical service data prediction and meta-learning algorithm, the problems of low efficiency and inability to implement active load management in the existing technology are solved, and efficient and active base station load balancing are achieved.
Patent Information
- Application Number
- CN202510396510.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing base station load balancing algorithm is inefficient in the face of rapid business changes and emergencies, and cannot implement active load management, resulting in network congestion and user service interruption.
By predicting the network service demand distribution characteristics based on historical service data, adjusting the base station resource distribution, and receiving the target service combination, and generating a load balancing strategy based on the target model. This method uses a meta-learning algorithm to quickly adjust the model parameters without adjusting the model structure to generate a load balancing strategy that conforms to the constraint model.
It realizes rapid adaptation and optimization of model parameters without changing the model structure, responding to rapid changes in business needs and sudden load situations, improving the efficiency and initiative of base station load balancing, and avoiding network congestion and user service interruption.
Smart Images

Figure CN120186680A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a base station load balancing method, apparatus, storage medium, and electronic device. Background Art
[0002] In modern communication networks, especially in the 5G and future 6G network environments, base station load balancing solutions have become a key technology to ensure network service quality and user experience. With the popularization of the mobile Internet and the rise of the Internet of Things, wireless networks are facing unprecedented challenges in service demand. In areas with dense population or frequent activities, such as large-scale events, concerts, and holiday shopping malls, the rapidly increasing user data traffic within a short period of time quickly overwhelms the carrying capacity of the base stations, resulting in network congestion and user service interruption.
[0003] The core of this problem lies in the limitations of existing base station load balancing algorithms. Traditional load balancing methods, such as rule-based static allocation strategies and dynamic allocation strategies that only perform passive adjustments when overload is detected, are no longer able to cope with rapidly changing and highly dynamic service scenarios. When the network environment or user behavior changes suddenly, these algorithms often need to be retrained or adjusted complexly, which not only consumes a large amount of time and computing resources, but may also cause service interruption during the adjustment process, reducing the user experience.
[0004] For the above problems, no effective solutions have been proposed yet. Summary of the Invention
[0005] This application provides a base station load balancing method, apparatus, storage medium, and electronic device to at least solve the technical problem that in the prior art, when the base station load balancing solution faces rapid service changes and emergencies, the traditional load balancing algorithm is inefficient and cannot achieve proactive load management.
[0006] According to one aspect of the present application, a base station load balancing method is provided, including: predicting the distribution characteristics of network service demands in a target scenario within a target time period based on historical service data; adjusting the resource distribution of the base station according to the distribution characteristics; receiving a target service combination within the target time period, where the target service combination includes M services and the priority of each service, and M is an integer greater than or equal to 1; based on the resource distribution of the base station, performing an adjustment operation on a target model according to the target service combination and a target algorithm, and generating a load balancing strategy for the target service combination based on the adjusted target model, where the target algorithm is used to adjust the model parameters of the target model based on the target service combination without adjusting the structure of the target model, so that the target model generates a load balancing strategy that conforms to a constraint model, where the constraint model at least includes a policy target, and the policy target is used to characterize that the service rate of services with a priority less than a preset value among the M services is greater than a preset rate, and the service time variance of the target service combination is less than a preset threshold.
[0007] Optionally, the constraint model is obtained through the following steps: setting an objective function, where the objective function is used to quantify the policy target; determining N constraint conditions according to a first model, where N is an integer greater than or equal to 1, and the first model is used to characterize the performance indicators of the M services; determining the constraint model according to the objective function and the N constraint conditions.
[0008] Optionally, the target model is obtained through the following steps: obtaining S historical service combinations, where each historical service combination includes multiple historical services, and each historical service includes a corresponding priority; performing multiple target operations on an initial model based on the S historical service combinations until the number of target operations is greater than or equal to a first iteration number, and obtaining the target model, where each target operation is used to train the initial model according to each historical service combination.
[0009] Optionally, the target operation includes the following steps: performing multiple target trainings on the initial model according to the i-th historical service combination until a first parameter combination of the initial model is obtained, where i is an integer greater than or equal to 1 and less than or equal to S, and the first parameter combination is used to characterize the model parameters of the initial model under the condition of conforming to the constraint model; updating the initial model according to the first parameter combination to obtain a first model; based on the first model, determining a second parameter combination by using the target algorithm, where the second parameter combination is a guiding parameter determined by the target algorithm and is used to adjust the model parameters of the first model; determining the target model according to the second parameter combination and the first model.
[0010] Optionally, the target training includes the following steps: adjusting the initial parameters of the initial model based on the i-th historical service combination to obtain a third parameter combination; updating the initial model according to the third parameter combination to obtain a second model; training the i-th historical service combination according to the second model to generate a first policy; and determining whether the third parameter combination is the first parameter combination according to the first policy and the constraint model.
[0011] Optionally, determining whether the third parameter combination is the first parameter combination according to the first policy and the constraint model includes: processing the i-th historical service combination according to the first policy to obtain a processing result, where the processing result at least includes the performance metrics of completing the i-th historical service combination; determining whether the processing result conforms to the constraint model; and determining that the third parameter combination is the first parameter combination of the initial model when the processing result conforms to the constraint model.
[0012] Optionally, after performing an adjustment operation on the target model based on the target service combination and generating a load balancing policy for the target service combination based on the adjusted target model, the method further includes: allocating resources to the target service combination according to the load balancing policy to obtain an allocation result; executing the target service combination based on the allocation result to obtain an execution result, where the execution result at least includes the performance metrics of completing the target service combination; and updating the target model according to the execution result.
[0013] According to another aspect of the present application, there is also provided a base station load balancing device, including: a prediction unit for predicting the distribution characteristics of the network service demand in the target time period for the target scenario based on historical service data; an adjustment unit for adjusting the resource distribution of the base station according to the distribution characteristics; a receiving unit for receiving the target service combination in the target time period, where the target service combination includes M services and the priority of each service, where M is an integer greater than or equal to 1; a generating unit for performing an adjustment operation on the target model based on the resource distribution of the base station, according to the target service combination and the target algorithm, and generating a load balancing policy for the target service combination based on the adjusted target model, where the target algorithm is used to adjust the model parameters of the target model based on the target service combination without adjusting the structure of the target model, so that the target model generates a load balancing policy that conforms to the constraint model, where the constraint model at least includes a policy target, where the policy target is used to characterize that the service rate of the service with a priority less than a preset value among the M services is greater than a preset rate, and the service time variance of the target service combination is less than a preset threshold.
[0014] According to another aspect of the present application, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned base station load balancing method.
[0015] According to another aspect of the present application, an electronic device is further provided, including one or more processors and a memory, where the memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned base station load balancing method.
[0016] In the present application, first, based on historical service data, the distribution characteristics of the network service demand in the target scenario within the target time period are predicted. Then, the resource distribution of the base station is adjusted according to the distribution characteristics. Next, the target service combination within the target time period is received, where the target service combination includes M services and the priority of each service, and M is an integer greater than or equal to 1. Finally, based on the resource distribution of the base station, the target model is adjusted based on the target service combination and the target algorithm, and a load balancing strategy for the target service combination is generated based on the adjusted target model. The target algorithm is used to adjust the model parameters of the target model based on the target service combination without adjusting the structure of the target model, so that the target model generates a load balancing strategy that conforms to the constraint model. The constraint model at least includes a policy target, where the policy target is used to characterize that the service rate of the service with a priority less than the preset value among the M services is greater than the preset rate, and the service time variance of the target service combination is less than the preset threshold. That is, through a combination of initially adjusting the resource distribution of the base station based on historical service data prediction and dynamic parameter adjustment, the purpose of quickly adapting and optimizing the model parameters to cope with the rapid change of business requirements and sudden load conditions is achieved without changing the model structure, thereby realizing the technical effect of efficient and proactive base station load balancing, and further solving the technical problems that in the prior art, the base station load balancing scheme has low efficiency of traditional load balancing algorithms and cannot achieve proactive load management when facing rapid business changes and sudden situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0018] Figure 1 is a flowchart of an optional base station load balancing method according to an embodiment of the present application;
[0019] Figure 2 is a schematic diagram of an optional base station load balancing device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To enable those skilled in the art to better understand the solution of this application, the technical solution in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0021] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0022] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. And the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure and application, etc., all comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set between this system and relevant users or institutions to provide corresponding operation entrances for users to choose to agree or refuse the automated decision-making results; if the user chooses to refuse, the expert decision-making process will be entered.
[0023] According to the embodiments of this application, a method embodiment of a base station load balancing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described here can be executed in a different order than here.
[0024] It should be noted that an intelligent processing system can be used as the execution subject of the base station load balancing method in the embodiments of this application. It can be understood that the base station load balancing method provided in the embodiments of this application can also be executed by other systems or devices as the execution subject, and the embodiments of this application do not make specific limitations on this.
[0025] Figure 1 It is a flowchart of an optional base station load balancing method according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0026] Step S101, predicting the distribution characteristics of network service demands in a target scenario within a target time period based on historical service data.
[0027] Optionally, historical service data refers to the usage records of different types of services in the network over a past period of time, including the request frequency, data volume, SLA level (Service Level Agreement, which refers to the level of service quality standards agreed between service providers and users), and actual network performance data (such as latency, packet loss rate, transmission rate, etc.) of each service.
[0028] Optionally, in the field of communication networks, especially for advanced application scenarios such as ultra-reliable low-latency communication, the SLA level defines various performance indicators of the service, including but not limited to delay time, packet loss rate, bandwidth, service availability, etc., and the specific numerical requirements for these indicators. Different levels of SLA correspond to different levels of service quality commitments.
[0029] Optionally, the distribution characteristics involve the statistical characteristics of service requests, such as the distribution of request frequencies, the distribution of packet sizes, the distribution of service priorities, etc. These characteristics can help predict future service demand patterns.
[0030] Optionally, the intelligent processing system analyzes historical service data to predict the distribution of network service demands in a future target time period (such as within the next few hours) in a specific scenario, including key information such as the request frequencies and data volumes of different types of services. This step is the basis for proactive load balancing, enabling the base station to prepare resources before the demand peak arrives.
[0031] Step S102, adjusting the resource distribution of the base station according to the distribution characteristics.
[0032] Optionally, resource distribution refers to how the base station allocates its limited resources (such as spectrum, bandwidth, transmission power, etc.) to different services to meet service demands and SLA levels.
[0033] Optionally, based on the predicted service demand distribution in the above steps, the base station adjusts its resource allocation strategy, giving priority to the demands of high-priority services to ensure the effective utilization of resources and leaving room for possible traffic surges.
[0034] Optionally, based on the prediction result of the traffic volume distribution, the intelligent processing system can predict the traffic hotspots in advance, enabling the base stations in the hotspots to actively perform load balancing and improving the average system throughput.
[0035] Step S103: Receive the target service combination within the target time period.
[0036] In step S103, the target service combination includes M services and the priority of each service.
[0037] In step S103, M is an integer greater than or equal to 1.
[0038] Optionally, the target service combination refers to the set of all services that the base station needs to process within the target time period and their priorities (i.e., SLA levels), where M represents the number of services.
[0039] Optionally, the intelligent processing system receives and processes service requests in real time at the current moment. These requests form the target service combination, including different service types and their respective priorities. This is the key to achieving dynamic load balancing, ensuring that the base station can respond to network changes in a timely manner and optimize resource allocation.
[0040] Step S104: Based on the resource distribution of the base station, adjust the target model according to the target service combination and the target algorithm, and generate a load balancing strategy for the target service combination based on the adjusted target model.
[0041] In step S104, the target algorithm is used to adjust the model parameters of the target model based on the target service combination without adjusting the structure of the target model, so that the target model generates a load balancing strategy that meets the constraint model.
[0042] In step S104, the constraint model includes at least a policy target.
[0043] In step S104, the policy target is used to represent that the service rate of the services with priorities less than the preset value among the M services is greater than the preset rate, and the service time variance of the target service combination is less than the preset threshold.
[0044] Optionally, the target algorithm refers to a specific algorithm for adjusting the target model parameters. In this embodiment, the target algorithm is a meta-learning algorithm. It can quickly adjust the model parameters according to the target service combination without changing the model structure to generate a load balancing strategy that meets the constraint model.
[0045] It should be noted that the meta - learning algorithm is an advanced machine - learning technology. Its core idea is to train a model on a series of related but different tasks so that the model can quickly adapt to new tasks. Specifically in this embodiment, the meta - learning algorithm is designed to solve the load - balancing problem among internal services of a base station, especially in scenarios facing rapid business changes and requiring proactive management. Meta - learning, also known as learning to learn, can be roughly divided into two methods: optimization - based and probability - based. Optimization - based meta - learning methods, such as MAML (Model - Agnostic Meta - Learning) and gradient - based meta - learning algorithms, find a set of parameters such that these parameters can achieve good performance on new tasks after a small number of gradient updates. Probability - based meta - learning methods, such as Bayesian Optimization and Variational Inference, construct a prior distribution of parameters among tasks and then perform posterior inference on new tasks to quickly obtain effective parameter configurations.
[0046] Optionally, the constraint model defines the conditions that the load - balancing strategy must meet. For example, the policy objective requires that the service rate of services with a priority lower than a certain preset value among M services must be greater than the preset rate, and at the same time, the variance of the service time of the entire service portfolio should be less than the preset threshold to ensure service quality and load balancing.
[0047] Optionally, the target model refers to the model that is specifically used to generate the load - balancing strategy after the parameters of the meta - learning algorithm are adjusted. It comprehensively considers historical data prediction, the current situation of resource distribution, service portfolio, and priority information, and can provide a customized solution for the real - time network environment.
[0048] Optionally, the intelligent processing system quickly adjusts the target model through the target algorithm (meta - learning) so that it can quickly adapt to changes in the target service portfolio and generate a load - balancing strategy that meets the requirements of the constraint model. After adjusting the parameters, the target model can intelligently improve the service rate of low - priority services while maintaining the service quality of high - priority services, and at the same time control the variance of service time to avoid excessive service - time fluctuations from damaging the user experience. Through this dynamic and intelligent adjustment process, the base station can achieve efficient and fair resource allocation and maintain good service quality even in complex scenarios with rapid changes in network demand.
[0049] Optionally, the intelligent processing system uses meta - learning to adjust the parameters based on the trained model, which speeds up the model training speed and significantly reduces costs.
[0050] As can be seen from the content of steps S101 to S104, in this application, first, the distribution characteristics of the network service demand of the target scenario in the target time period are predicted based on historical service data. Then, the resource distribution of the base station is adjusted according to the distribution characteristics. Then, the target service combination in the target time period is received, where the target service combination includes M services and the priority of each service, where M is an integer greater than or equal to 1. Finally, based on the resource distribution of the base station, the target model is adjusted based on the target service combination and the target algorithm, and a load balancing strategy for the target service combination is generated based on the adjusted target model. The target algorithm is used to adjust the model parameters of the target model based on the target service combination without adjusting the structure of the target model, so that the target model generates a load balancing strategy that meets the constraint model. The constraint model at least includes a policy objective, where the policy objective is used to characterize that the service rate of the service with a priority less than a preset value among the M services is greater than a preset rate, and the service time variance of the target service combination is less than a preset threshold. That is, through the combination of predicting based on historical service data, initially adjusting the resource distribution of the base station, and dynamic parameter adjustment, the purpose of quickly adapting to and optimizing model parameters to cope with rapid changes in business requirements and sudden load situations is achieved without changing the model structure, thereby achieving the technical effect of efficient and proactive base station load balancing, and further solving the technical problems in the prior art that in the face of rapid changes in business and sudden situations, the traditional load balancing algorithm is inefficient and cannot achieve proactive load management.
[0051] In an alternative embodiment, the intelligent processing system first sets a target function, where the target function is used to quantify the policy objective. Then, N constraint conditions are determined according to the first model, where N is an integer greater than or equal to 1, and the first model is used to characterize the performance indicators of M services. Finally, the constraint model is determined according to the target function and the N constraint conditions.
[0052] Optionally, the intelligent processing system defines a target function, and the purpose of this function is to quantify the load balancing policy objectives of different service combinations inside the base station. The specific objectives include maximizing the service rate of the service with the lowest service level and minimizing the service time variance at the same time to ensure the QoS (Quality of Service) of all services.
[0053] Optionally, the first model refers to the URLLC model (Ultra-Reliable and Low Latency Communications model), which is a mathematical model used to describe the characteristics and behaviors of ultra-reliable and low-latency communication services. In this embodiment, the URLLC model is not only used to describe high-priority services, but also used to characterize the performance metrics and mutual influences of all M services inside the base station.
[0054] Optionally, the intelligent processing system sets N constraint conditions based on the URLLC model, where N represents the number of constraints and is an integer greater than or equal to 1. These constraint conditions may include the minimum service rate, the maximum latency, the minimum requirement for packet reliability, etc., to ensure that the SLA level of the service can be met under any load balancing strategy.
[0055] Optionally, according to the set objective function and N constraint conditions, the intelligent processing system constructs a constraint model. The constraint model is a mathematical framework for a multi-objective optimization problem, which combines the objective function and the constraint conditions and is used to guide the training process of the meta-learning algorithm to ensure that the finally generated load balancing strategy can not only maximize the objective function, but also meet all preset SLA requirements and performance metrics.
[0056] Optionally, this embodiment models the downlink transmission scenario in a heterogeneous network. All users served by the base station adopt the URLLC model. Inside the base station, services are classified into SLA levels according to latency, packet reliability, and downlink user rate. Different services have different SLA service levels, thus generating different service combination situations.
[0057] Optionally, the base station set is represented as The available radio resources are divided into N radio resources in the time-frequency domain. Each radio resource has a bandwidth B in the frequency domain and each transmission time interval lasts 1 ms in the time domain. Therefore, there are a total of N available radio resources in a time slot, and the available time slots can be further divided into k smaller units.
[0058] Optionally, an SLA is a formal agreement between a service provider and a tenant or between service providers, based on which the provided service levels are clearly defined. Each SLA contains a specific number of elements, which are called metrics, used to describe the level and quantity of communication services and measure the performance characteristics of the service object. In this embodiment, different standards are defined for different URLLC services according to the network SLA metric capability levels. Considering the load balancing of service combinations with different levels, such as the service combination of SLA1 and SLA2. Among them, the larger the number of the capability level, the higher the priority of the service. This embodiment stipulates that when the service traffic with a higher SLA metric capability level arrives, the base station will have a probability of delaying the ongoing transmission of the service with a lower level.
[0059] Optionally, in this embodiment, the users served by the base station uniformly adopt the URLLC model. First, use the symbol u s to represent the URLLC service with the SLA level s. Since services with different SLA levels have different priorities, transmitting URLLC traffic with a higher SLA level will affect the rate of the ongoing service with a lower level. Therefore, the decision variable is introduced to represent the service decision of the w-th URLLC service at the current moment, as shown in formula (1):
[0060]
[0061] where b represents the base station serial number, n represents the current radio resource serial number, and k represents the current time slot.
[0062] Optionally, define the signal-to-noise interference ratio of the w-th URLLC service as shown in formula (2):
[0063]
[0064] where and respectively represent the transmission power and channel gain of the w-th URLLC service on the n-th radio resource with the base station serial number b at the current moment, and respectively represent the transmission power and channel gain of the w-th URLLC service on the n-th radio resource with the base station serial number not b at the current moment, and σ 2 is the noise power.
[0065] Optionally, to avoid transmission delay, the block length of the URLLC service should be limited. Define the achievable data rate of the URLLC service The formula is shown in (3):
[0066]
[0067] wherein, represents the number of symbols per micro-slot, and Q -1 (x) represents the Gaussian inverse cumulative distribution function, represents the channel dispersion, which determines the channel randomness of the service and can be expressed by Equation (4):
[0068]
[0069] Optionally, since different services have different priorities, when a service with a higher service level arrives, there is a probability that a service with a lower service level will be postponed, which will affect the system capacity and reliability. Therefore, considering the service rates and service times of different services, this embodiment establishes a multi-objective optimization problem (policy objective) to maximize the service rate of the service with the lowest service level while minimizing the variance of the service time, and needs to satisfy the constraints such as the delay limit, reliability limit, and downlink transmission rate of the corresponding service. For URLLC services, it is assumed that users will create small data packet fragments, and the micro-slots of the data packets within the time slot t follow a Poisson point process distribution. Use the random variable ψ k (t) to represent the number of arriving data packets, as shown in Equation (5):
[0070]
[0071] wherein, ψ(t) represents the total amount of URLLC data packets arriving within the time slot t. Based on this, the reliability of the URLLC service can be obtained from Equation (6):
[0072]
[0073] wherein, κ represents the data packet size of the URLLC service. Equation (6) shows that the outage probability of the URLLC service should not exceed the threshold of the corresponding service level. Therefore, the constraint model of multiple URLLC service combinations can be expressed by the following mathematical formula:
[0074]
[0075] wherein, σ(t) in Equation (7) represents the variance of the service time of different services, and the overall objective is to maximize the service rate of the service with the lowest service level while minimizing the variance of the service time; Equation (8) means that at any time t, for any user equipment n and base station b, only one service k can be in progress on any given radio resource w; Equation (9) illustrates the decision variable The value of can only be 0 or 1, which reflects the binary nature of resource allocation: resources are either used or not used, with no intermediate state; Equation (10) guarantees the reliability of URLLC and stipulates the upper limit of the probability of service interruption; Equation (11) is the transmission rate constraint for the current service level, ensuring the service rate for base station b and user equipment n on service combination (s, w) must at least reach or exceed the minimum service rate SLA stipulated by the service level agreement SLA s .
[0076] As can be seen from the above, the intelligent processing system can achieve active load balancing between services inside the base station in a systematic and mathematical way through the above method. The objective function ensures that the SLA levels of all services are optimally considered, while the N constraints provide the minimum threshold for the quality of service of the system, including but not limited to the minimum requirements for service rate and the limitation of service time variance. The establishment of the constraint model enables the intelligent processing system to quickly generate and execute load balancing strategies that meet the SLA levels and service performance indicators when facing real-time changing service combinations. This method not only improves the transmission efficiency of high-priority services but also takes into account the basic performance requirements of low-priority services, thus realizing the efficient utilization of network resources and the improvement of the overall service quality, enhancing the adaptability and robustness of the system. In a rapidly changing mobile network environment, this method can significantly improve the user experience, reduce the service interruption rate and latency, and optimize the spectrum efficiency and network resource allocation.
[0077] In an optional embodiment, the target model is obtained through the following steps: The intelligent processing system first obtains S historical service combinations, where each historical service combination includes multiple historical services, and each historical service includes a corresponding priority. Then, based on the S historical service combinations, multiple target operations are performed on the initial model until the number of target operations is greater than or equal to the first iteration number, and the target model is obtained, where each target operation is used to train the initial model according to each historical service combination.
[0078] Optionally, the historical service combination refers to the historical service request records collected by the intelligent processing system, and each combination contains different service types and quantities requested within a specific time window. The historical service is a single service instance that makes up the historical service combination, and each service instance is marked with its service type and priority information.
[0079] Optionally, the initial model: refers to a model that has not been trained with historical service data. Target operation: Each operation involves using a historical service combination as input to train the initial model to adjust the model parameters so that it can better predict and process similar service combinations. Target model: After multiple rounds of target operations training, a model with optimized performance and adjusted parameters, which is used to generate a load balancing policy that meets the requirements of the constraint model.
[0080] Optionally, the intelligent processing system extracts S historical service combinations from the historical database. These combinations record the specific situations of service requests in the base station at different times and scenarios in the past, including key information such as service type, service volume, and service priority. Here, S is the number of historical service combinations, and its value is greater than or equal to 1. The more historical service combinations, the more experience the system can learn and extract from richer historical scenarios.
[0081] Optionally, the intelligent processing system uses the obtained S historical service combinations to perform multiple rounds of training on the initial model, that is, the target operation. Each target operation includes using a historical service combination to train the model once. Through continuous iteration, the model gradually learns how to make optimal load balancing decisions when facing different service priorities and combinations. The first number of iterations is a preset number of training rounds. When the number of target operations reaches or exceeds this preset value or after all S historical service combinations have been trained, the training process ends. The system will obtain a target model. When dealing with the load balancing problem between internal services in the base station, this model can more accurately predict and respond to service demands while considering the priorities of different services.
[0082] As can be seen from the above, iterative training of the model based on historical service data enables the intelligent processing system to generate a highly adaptable and intelligent load balancing target model. The training process of the target model allows the system to fully learn and understand the service demand patterns and priority changes in historical service combinations, so that when facing new service combinations, it can quickly make decisions and generate load balancing policies. The training process based on S historical service combinations also enhances the robustness of the model, enabling it to handle various complex service combination scenarios and avoiding performance limitations caused by a single training scenario of the model.
[0083] In an alternative embodiment, the target operation comprises the following steps: The intelligent processing system first performs multiple target trainings on the initial model according to the i-th historical service combination until a first parameter combination of the initial model is obtained, where i is an integer greater than or equal to 1 and less than or equal to S, and the first parameter combination is used to characterize the model parameters of the initial model under the constraint model; update the initial model according to the first parameter combination to obtain a first model, and then based on the first model, use the target algorithm to determine a second parameter combination, where the second parameter combination is a guiding parameter determined by the target algorithm and is used to adjust the model parameters of the first model, and finally determine the target model according to the second parameter combination and the first model.
[0084] Optionally, the first parameter combination refers to the set of model parameters that the initial model can meet the requirements of the constraint model after multiple trainings. These parameters are the core of the model and determine the behavior and performance of the model in a specific scenario.
[0085] Optionally, the intelligent processing system reads and analyzes the data of the i-th historical service combination. These data may include past service request frequencies, packet sizes, service priorities, and actual resource allocation situations, etc. Then, the system uses these data to train the initial model, adjusts the model parameters through the backpropagation algorithm to minimize the prediction error and optimize the objective function. This training process may involve thousands or millions of iterations until the model parameters are stable and meet the requirements of the constraint model, thereby obtaining the first parameter combination.
[0086] Optionally, the intelligent processing system uses the first parameter combination to update the parameters of the initial model, thereby obtaining a first model that is more in line with the characteristics of the historical service combination. This update process may involve weight adjustment, bias correction, etc., to ensure that the model can exhibit better prediction and decision-making capabilities while maintaining the original structure. Then the intelligent processing system uses the target algorithm to perform a meta-learning process based on the first model. This process involves rapid iteration, adjusts the hyperparameters of the model according to the characteristics of the current target service combination and the requirements of the constraint model, and obtains the second parameter combination. Through this adjustment, the first model will better adapt to the current scenario and generate a more accurate load balancing strategy. Finally, the intelligent processing system uses the second parameter combination to update the parameters of the first model to generate the target model. This model will be used to process the target service combination in real time and dynamically adjust the resource allocation according to the requirements of the constraint model to generate the optimal load balancing strategy.
[0087] As can be seen from the above, through the training of historical service combinations and the rapid parameter tuning of the meta-learning algorithm, the intelligent processing system can quickly generate the optimal load balancing strategy that meets the requirements of the constraint model. This method improves the generalization ability and adaptability of the model, enabling the system to maintain efficient and fair resource allocation even under rapidly changing service demands and network environments, meet the SLA levels of all services, and enhance the overall network performance and user experience. Specifically, the first parameter combination obtained through training with historical data ensures the stability and effectiveness of the model in historical scenarios; while the second parameter combination obtained through the meta-learning algorithm enables the model to quickly converge in new scenarios and generate strategies that conform to the constraint model. The entire process not only improves the training efficiency of the model but also enhances the application flexibility of the model under different service combinations, realizes the dynamic optimization of the load balancing strategy, and thus significantly improves the utilization efficiency of network resources and the overall performance of the system.
[0088] In an alternative embodiment, the target training includes the following steps: First, the intelligent processing system adjusts the initial parameters of the initial model based on the i-th historical service combination to obtain a third parameter combination. Then, it updates the initial model according to the third parameter combination to obtain a second model. Next, it trains the i-th historical service combination based on the second model to generate a first strategy. Finally, it determines whether the third parameter combination is the first parameter combination according to the first strategy and the constraint model.
[0089] Optionally, the intelligent processing system reads the data of the i-th historical service combination, including information such as the type, frequency, SLA level, and service priority of the service request. Then, the system uses this information to train the initial model, adjusts the initial parameters of the model through machine learning techniques such as backpropagation, and obtains a third parameter combination that can better predict and process the historical service combination. Next, the intelligent processing system uses the third parameter combination to update the parameters of the initial model and generate a second model. This step involves re-initializing certain layers of the model with the third parameter combination or directly using them to update the corresponding model weights. The updated second model will be used in the subsequent training process to generate a more effective load balancing strategy. Then, the intelligent processing system uses the second model to perform a deep learning or reinforcement learning training process on the i-th historical service combination. During the training process, the model will attempt to predict and process different service types and priorities in the service combination to minimize the service time variance while maximizing the rate of low-priority services and meeting the policy objectives in the constraint model. After the training is completed, the system will obtain a first policy, which is the specific performance of the second model on the i-th historical service combination and is used to guide the resource allocation within the base station. Finally, the intelligent processing system compares the first policy with the constraint model to check whether the service rates and service time variances of all services are within the constraint range, especially whether the service rate of services with a lower SLA level is greater than the preset value and whether the service time variance is less than the preset threshold. If the first policy meets all the requirements of the constraint model, the third parameter combination will be confirmed as the first parameter combination for the next model update; otherwise, the system will adjust the third parameter combination again based on the feedback information until the generated policy fully meets the constraint model.
[0090] As can be seen from the above, through staged model training and parameter adjustment, it is ensured that the intelligent processing system can automatically generate the optimal load balancing strategy that meets the requirements of the constraint model. This method not only improves the efficiency of model training, but also enhances the accuracy of the strategy, can effectively handle different service combination scenarios, guarantee the SLA level of all services, especially outstanding in the rate optimization of low-priority services and the overall system delay control. The overall effects are reflected in the following points: (1) Improve the training speed and quality of the model. Through gradual parameter adjustment and strategy optimization, the model can find the first parameter combination that meets the conditions within fewer training rounds, reducing the training time and resource consumption; (2) Achieve the dynamic adaptability of the strategy. Through the training of different historical service combinations, the model can learn and understand the change patterns of service requirements, generate strategies that can quickly adapt to new service combinations, and improve the flexibility and response speed of the system; (3) Guarantee the SLA level. Through the setting of the constraint model, it is ensured that the strategies generated by the model can meet the SLA requirements of all services, including key indicators such as service rate, delay and packet reliability, thus improving user satisfaction and network service quality.
[0091] In an alternative embodiment, the intelligent processing system first processes the i-th historical service combination according to the first strategy to obtain a processing result, where the processing result at least includes the performance indicators of completing the i-th historical service combination, and then determines whether the processing result meets the constraint model. When the processing result meets the constraint model, it is determined that the third parameter combination is the first parameter combination of the initial model.
[0092] Optionally, based on the first policy, the intelligent processing system processes the i-th historical service combination, simulating the entire process of service request arrival, resource allocation, data transmission, and service completion. This process may involve operations such as allocating service requests to the wireless resources of different base stations, adjusting the transmission power, allocating channel bandwidth, and controlling the service time variance. After the processing is completed, the system records and analyzes the obtained processing results, including performance indicators such as the service completion time, rate, packet reliability, etc., and whether abnormal situations such as service interruption and timeout occur. Then the intelligent processing system compares the processing results obtained in the previous step with the constraint model to check whether the processing results meet the various indicators and constraint conditions specified in the constraint model. If the processing results meet the constraint model, it indicates that the first policy performs well in processing the i-th historical service combination and can achieve the expected service quality and system performance goals. Once the system confirms that the processing results conform to the constraint model, it will immediately perform the model parameter update operation. This usually involves assigning the third parameter combination to the parameters of the initial model to make it the first parameter combination, thereby updating the model state and preparing for the processing and policy generation of subsequent service combinations. The updated model will become the second model, which is used to process more complex service combinations to achieve more efficient and intelligent load balancing.
[0093] Optionally, the overall framework of the target model includes two parts: inner-layer training and outer-layer training. Among them, the inner-layer training introduces a reinforcement learning algorithm, aiming to train a highly adaptable load balancing mechanism for different business combination scenarios. Through reinforcement learning, the inner-layer model can dynamically adjust the policy according to the specific business combination, thereby achieving efficient load resource allocation. In the outer-layer training, a meta-learning algorithm is used to guide the inner-layer training. The core role of meta-learning is to quickly adapt to different service combination scenarios by only adjusting the hyperparameters of the model while keeping the model type and structure unchanged. This method improves the convergence speed of the model in new scenarios, enabling the model to complete optimization in a short time. This framework combines the advantages of reinforcement learning and meta-learning, ensuring both the flexibility of the inner-layer model in specific tasks and the efficient inter-task transfer through outer-layer meta-learning, providing strong support for complex and changing service combination scenarios.
[0094] Optionally, since the service has multiple SLA levels, in this embodiment, different services are regarded as multiple agents, and an algorithm based on reinforcement learning is proposed to train the load balancing model. The reinforcement learning task is usually described using the Markov Decision Process (MDP). The MDP consists of a tuple (S, A, p, R, γ), where S is the state, A is the action, p is the state transition probability, R is the reward, and γ is the discount factor. After giving a policy π, first define the cumulative return G t As shown in formula (12):
[0095]
[0096] where G t is the cumulative reward after taking an action at time step t, i.e., the sum of rewards obtained by the agent at all future time steps; R t+1 represents the next time step t + 1 after taking an action, R t+2 represents the second time step t + 2, and γ is a discount factor between 0 and 1; is an expression in the form of an infinite series that summarizes the rewards R obtained by the agent at each future time step t + k + 1 starting from time step t t+k+1 and each term is weighted by the corresponding power of the discount factor γ k weighted.
[0097] Optionally, the mathematical expectation of the cumulative reward is used to define the value of the current state, which is called the value function. Use v π (s) to represent the value function of state s under policy π, as shown in formula (13):
[0098]
[0099] where S t = s represents the condition that the current state is s. The value function v π (s) for state s is the sum of the expected future returns in this state.
[0100] Optionally, Soft Actor-Critic (SAC) is an algorithm in reinforcement learning that combines policy gradient and value function estimation, applicable to solving discrete and continuous action space problems, and is an off-policy algorithm. SAC realizes the decision-making and learning process through two main components, an actor and a critic. The actor is responsible for generating actions, and the critic evaluates the value of these actions. SAC combines the idea of maximum entropy reinforcement learning, aiming to increase the policy entropy (i.e., exploration) while maximizing the cumulative reward, encouraging the agent to be more random when choosing actions and being more adaptable in the face of interference.
[0101] Optionally, in the reinforcement learning algorithm used in this embodiment, the agent represents each service, and the definitions of its state set, action set, reward, and environment are as follows. The state space S of a single agent contains the following information: the latency of the current service, the packet size of the current service, the number of packets to be transmitted at the current moment, and the service SLA level. The states of all agents can be combined into a global state vector S global, which is used to describe the dynamic changes of the entire system. The traffic prediction results based on ensemble learning will be used as part of the state space and affect the number of data packets.
[0102] Optionally, the action a of a single agent represents whether data transmission is performed for this type of service at the current moment. It can be expressed by the formula as shown in (14):
[0103]
[0104] Optionally, since there are N types of services in total and each service has two choices, the size of the action space A is 2^N.
[0105] Optionally, the design goal of the reward function R is to optimize the load balancing performance of the overall system while meeting the SLA requirements of different services. The specific definition is shown in formula (15):
[0106]
[0107] where is the service rate of the service with the lowest service level among all services, and σ(t) is the variance of the service times of multiple services. By maximizing R, it is possible to minimize the overall system delay while ensuring the basic performance of low-priority services.
[0108] Optionally, the state of the environment changes dynamically over time and is affected by the request arrival rates of different services, the amount of data transmitted by the current service, etc. The core functions of the environment include: 1) receiving the action set A of the agent and updating the system state S global ; 2) calculating the reward R according to the new state; 3) returning the state S t+1 and the reward R t .
[0109] Optionally, in meta-reinforcement learning, the agent tries to learn a policy with parameter θ, which can solve multiple tasks in a given distribution . Each task is an MDP process, which consists of its input s i , output a i , loss function L i , transition function P i , reward function R i and episode length H i . Generally, the meta-reinforcement learning method has two steps, namely the meta-training process and the meta-testing process. During the meta-training process, the agent is trained on a set of meta-training tasks , and during the meta-testing process, the agent is evaluated on a set of test tasks . It is usually assumed that the training set and the test set come from the same distribution but can be different from each task has a training set and a validation set For each task, the goal is to start from the parameter θ and use the training set to learn the task-specific parameters such that the loss L on the validation set i is minimized. Alg(.) represents the algorithm for updating the task-specific parameters.
[0110] Optionally, in this embodiment, we focus on gradient-based meta-reinforcement learning methods, such as MAML (Model-Agnostic Meta-Learning). The meta-training phase of such methods can be regarded as a two-layer optimization problem, whose goal is to learn the optimal meta-parameter θ such that:
[0111]
[0112] where Equation (16) represents finding a set of parameters that minimize the function F(θ), and F(θ) is represented as shown in (17):
[0113]
[0114] Optionally, in MAML, the inner-layer optimization is solved by the following one (or more) gradient descent steps:
[0115]
[0116] where θ i represents the model parameters for a specific task ; This is the formula expression of the gradient descent update rule, which shows how to adjust the task-specific parameters θ from the meta-parameter θ i ; θ is the initial parameter; β is the step size of the inner-layer optimization; is the gradient of the loss function L i with respect to the parameter θ, which points in the direction of the fastest increase in the loss function value.
[0117] Optionally, for the multi-objective load balancing problem of this embodiment, each task is associated with a specific weight vector ω iThe corresponding Markov decision process (which means that each load balancing task, such as resource allocation under different service combinations, can be regarded as an MDP with state S, action a, reward r, and state transition probability p). The learning problem represented by formula (18) is solved by alternately executing two optimization steps: 1) Task adaptation (inner optimization): Starting from the meta-policy parameters θ, multiple task-specific policies are learned (at this stage, the algorithm uses the pre-trained meta-policy parameters θ as a starting point, and according to the characteristics and feedback (reward) of each task, adjusts the policy parameters through inner optimization to form specific policies for each task , which is equivalent to performing local policy optimization on each task to better serve the goals of that task); 2) Meta-adaptation (outer optimization): Adjust the meta-parameters using the trajectories sampled from the adapted policies (after inner optimization, for each task, a set of adapted policy parameters is obtained. Next, the algorithm will use these policies to sample a series of trajectories (state-action sequences) in the environment and update the meta-policy parameters θ through the feedback (such as rewards) of these trajectories. This process is achieved by calculating the policy gradient or other optimization methods, aiming to make the meta-policy parameters better initialize the learning of future tasks). These two steps are repeated for a fixed number of meta-iterations N meta . Once the training is completed, the obtained meta-policy can be used as the initial parameters to quickly learn the optimal solution for new tasks. This means that when encountering an unseen load balancing task, the model does not have to start training from scratch, but can start from the pre-trained meta-policy parameters and only requires a small number of updates to achieve good performance, which greatly saves the adaptation cost and time for new tasks.
[0118] Optionally, the main process of the meta-learning-based load balancing algorithm (target model), specifically, it adopts the framework of MAML to achieve fast adaptation of load balancing policy learning for different service scenarios. The following details each step:
[0119] Step 1 Input: This part defines the input parameters required when the algorithm starts, including: preference distribution p(ω), number of meta-iterations N meta , number of tasks N in each meta-iteration task, the number of trajectories M sampled for each task; among them, the preference distribution represents the occurrence probabilities of different service combinations or business scenarios. In this algorithm, it is used to sample the weight vector for training on different tasks; the number of meta-iterations: indicates how many rounds of meta-learning iterations the algorithm will perform; the number of tasks in each meta-iteration: in each round of meta-iteration, the algorithm will process how many independent tasks, which come from different business combinations or scenarios; the number of trajectories sampled for each task: for each task, how many trajectory data generated after the agent executes actions will the algorithm sample for updating the policy parameters.
[0120] Step 2 Randomly initialize the meta-policy π θ : At the beginning of the algorithm, the parameters θ of the model need to be randomly initialized. The meta-policy π θ here refers to an initial policy function, and its parameters θ will be optimized through subsequent meta-learning processes.
[0121] Step 3 Loop through t = 0, 1, …, N meta :
[0122] Step 3.1 Sample N preference vectors ω i p(ω): Sample N task preference vectors from the preference distribution, and these vectors represent the characteristics of each task, such as service combinations or business scenarios;
[0123] Step 3.2 Loop through each weight vector ω i :
[0124] Step 3.2.1 Sample M trajectories according to π θ and denote them as Use the current meta-policy π θ to execute actions (M) times in the simulated load balancing environment to generate (M) pieces of trajectory data. The trajectory data contains information on the state, action, reward, and new state of the agent when executing actions in the environment, and will be used for parameter adjustment;
[0125] Step 3.2.2 Use to estimate Based on the sampled trajectory data to calculate the gradient of the policy on the current task where L i is the loss function for task i.
[0126] Step 3.2.3 Calculate the parameter Use the calculated gradient to update the policy parameter θ with the learning rate β to obtain the specific parameter θ for task i i .
[0127] Step 3.2.4 According to the policy Collection trajectory Use the updated parameter θ i to generate another set of verification trajectory data under the same business scenario This set of data will be used to evaluate the performance of the policy on the task and perform meta - learning updates.
[0128] Step 3.3 Update the policy After completing the inner - layer training and verification of all tasks, use the gradient information of all tasks to update the parameters θ of the meta - policy through the learning rate η of meta - learning. This is the key to meta - learning, which enables the model to learn shareable patterns from multiple tasks, thus accelerating the adaptation to new tasks.
[0129] Step 4 The algorithm ends. After completing N meta rounds of iteration, the algorithm ends. At this time, the meta - policy parameter θ should have been optimized to the extent that it can quickly adapt to new tasks.
[0130] As can be seen from the above, through actual policy application and result verification, it is ensured that the intelligent processing system can generate the optimal load - balancing policy that meets the constraint model. This method improves the practicality and performance reliability of the policy. Through the simulation and verification of historical service combinations, the system can accurately evaluate the effectiveness of the policy and ensure that its application in the real scenario will not violate the service - level agreement (SLA) and the usage constraints of network resources.
[0131] In an alternative embodiment, the intelligent processing system allocates resources to the target service combination according to the load - balancing policy to obtain the allocation result, and then executes the target service combination based on the allocation result to obtain the execution result. The execution result at least includes the performance metrics for completing the target service combination, and then the target model is updated according to the execution result.
[0132] Optionally, the intelligent processing system reads the load balancing policy generated by the target model and allocates resources to each service in the target service combination according to the guiding principles in the policy. For example, the system may preferentially allocate more wireless resources to high-priority services, while allocating resources to low-priority services on the premise of meeting their basic performance requirements. The allocation results will record in detail the resource type, quantity, and time allocated to each service, as well as the expected service performance metrics. Then the intelligent processing system executes the target service combination according to the allocation results and monitors the performance during the service execution in real time, including data transmission rate, latency, packet reliability, and service time variance, etc. After the service execution is completed, the system will collect the execution performance data of all services to form an execution result, which will be used to evaluate the effectiveness of the load balancing policy and feedback to the target model for learning and adjustment. Finally, the intelligent processing system analyzes the execution result, evaluates the difference between the actual service performance and the expected goal, and then uses machine learning techniques such as backpropagation to update the parameters of the target model using the execution result. This update process may involve multiple rounds of iteration until the model can generate a policy closer to the expected execution result in a similar scenario. The updated model is the latest target model, which can more accurately predict service requirements and generate effective resource allocation policies.
[0133] As can be seen from the above, the intelligent processing system realizes the active load balancing of the services inside the base station through a closed-loop policy generation, execution, and model update process, optimizes the resource allocation efficiency, and improves the service quality and user experience.
[0134] The embodiment of the present application also provides a base station load balancing device. It should be noted that the base station load balancing device of the embodiment of the present application can be used to execute the base station load balancing method provided by the embodiment of the present application. The base station load balancing device provided by the embodiment of the present application will be introduced below.
[0135] According to the embodiment of the present application, there is also provided a device for implementing the above base station load balancing method. Figure 2 It is a schematic diagram of an optional base station load balancing device according to the embodiment of the present application, as Figure 2 shown. The device includes: a prediction unit 201, an adjustment unit 202, a receiving unit 203, and a generating unit 204.
[0136] Optionally, a prediction unit 201 is configured to predict the distribution characteristics of the network service demand of the target scenario within the target time period based on historical service data; an adjustment unit 202 is configured to adjust the resource distribution of the base station according to the distribution characteristics; a receiving unit 203 is configured to receive a target service combination within the target time period, where the target service combination includes M services and the priority of each service, and M is an integer greater than or equal to 1; a generating unit 204 is configured to, based on the resource distribution of the base station, perform an adjustment operation on the target model according to the target service combination and a target algorithm, and generate a load balancing policy for the target service combination based on the adjusted target model, where the target algorithm is used to adjust the model parameters of the target model based on the target service combination without adjusting the structure of the target model, so that the target model generates a load balancing policy that conforms to a constraint model, where the constraint model at least includes a policy target, and the policy target is used to characterize that the service rate of the service with a priority less than a preset value among the M services is greater than a preset rate, and the service time variance of the target service combination is less than a preset threshold.
[0137] Optionally, the generating unit 204 includes: a first setting subunit, a first determining subunit, and a second determining subunit. The first setting subunit is configured to set an objective function, where the objective function is used to quantify the policy target; the first determining subunit is configured to determine N constraint conditions according to a first model, where N is an integer greater than or equal to 1, and the first model is used to characterize the performance indicators of the M services; the second determining subunit is configured to determine a constraint model according to the objective function and the N constraint conditions.
[0138] Optionally, the generating unit 204 includes: a first obtaining subunit and a first processing subunit. The first obtaining subunit is configured to obtain S historical service combinations, where each historical service combination includes a plurality of historical services, and each historical service includes a corresponding priority; the first processing subunit is configured to perform a plurality of target operations on an initial model based on the S historical service combinations until the number of target operations is greater than or equal to a first iteration number, and obtain a target model, where each target operation is used to train the initial model according to each historical service combination.
[0139] Optionally, the first processing subunit includes: a first training module, a first updating module, a first determining module, and a second determining module. Among them, the first training module is configured to perform multiple target trainings on the initial model according to the i-th historical service combination until a first parameter combination of the initial model is obtained, where i is an integer greater than or equal to 1 and less than or equal to S, and the first parameter combination is used to represent the model parameters of the initial model under the constraint model; the first updating module is configured to update the initial model according to the first parameter combination to obtain a first model; the first determining module is configured to determine a second parameter combination based on the first model by using a target algorithm, where the second parameter combination is a guiding parameter determined by the target algorithm and is used to adjust the model parameters of the first model; the second determining module is configured to determine a target model according to the second parameter combination and the first model.
[0140] Optionally, the first training module includes: a first adjustment sub-module, a first updating sub-module, a first generation sub-module, and a first determining sub-module. Among them, the first adjustment sub-module is configured to adjust the initial parameters of the initial model based on the i-th historical service combination to obtain a third parameter combination; the first updating sub-module is configured to update the initial model according to the third parameter combination to obtain a second model; the first generation sub-module is configured to train the i-th historical service combination according to the second model to generate a first policy; the first determining sub-module is configured to determine whether the third parameter combination is the first parameter combination according to the first policy and the constraint model.
[0141] Optionally, the first determining sub-module includes: a first processing component, a first judgment component, and a first determining component. Among them, the first processing component is configured to process the i-th historical service combination according to the first policy to obtain a processing result, where the processing result at least includes the performance index of completing the i-th historical service combination; the first judgment component is configured to judge whether the processing result conforms to the constraint model; the first determining component is configured to determine that the third parameter combination is the first parameter combination of the initial model when the processing result conforms to the constraint model.
[0142] Optionally, the base station load balancing device further includes: a distribution unit, an execution unit, and an update unit. Among them, the distribution unit is configured to perform resource allocation on the target service combination according to the load balancing policy to obtain a distribution result; the execution unit is configured to execute the target service combination based on the distribution result to obtain an execution result, where the execution result at least includes the performance index of completing the target service combination; the update unit is configured to update the target model according to the execution result.
[0143] According to another aspect of the present application, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned base station load balancing method.
[0144] According to another aspect of the present application, an electronic device is further provided, including one or more processors and a memory for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned base station load balancing method.
[0145] The serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0146] In the above-mentioned embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0147] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0148] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0149] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0150] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0151] The foregoing are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A base station load balancing method, characterized in that: include: Predict the distribution characteristics of network service demand in target scenarios within target time periods based on historical service data; adjusting the resource distribution of the base station according to the distribution characteristics; Receiving a target service combination within the target time period, wherein the target service combination includes M services and a priority of each of the services, wherein M is an integer greater than or equal to 1; Based on the resource distribution of the base station, the target model is adjusted according to the target service combination and the target algorithm, and a load balancing strategy for the target service combination is generated based on the adjusted target model, wherein the target algorithm is used to adjust the model parameters of the target model based on the target service combination without adjusting the structure of the target model, so that the target model generates a load balancing strategy that conforms to a constraint model, wherein the constraint model includes at least a policy target, wherein the policy target is used to characterize that a service rate of a service with a priority less than a preset value among the M services is greater than a preset rate, and that a service time variance of the target service combination is less than a preset threshold.
2. The base station load balancing method according to claim 1, characterized in that: The constraint model is obtained by the following steps: Setting an objective function, wherein the objective function is used to quantify the policy goal; Determining N constraint conditions according to a first model, where N is an integer greater than or equal to 1, and the first model is used to characterize the performance indicators of the M services; The constraint model is determined according to the objective function and the N constraint conditions.
3. The base station load balancing method according to claim 1, characterized in that: The target model is obtained by the following steps: Acquire S historical service combinations, wherein each of the historical service combinations includes multiple historical services, and each of the historical services includes a corresponding priority; The target model is obtained by performing multiple target operations on the initial model based on the S historical service combinations until the number of target operations is greater than or equal to the first iteration number, wherein each target operation is used to train the initial model according to each historical service combination.
4. The base station load balancing method according to claim 3, characterized in that: The target operation includes the following steps: Performing multiple target training on the initial model according to the i-th historical service combination until a first parameter combination of the initial model is obtained, wherein i is an integer greater than or equal to 1 and less than or equal to S, and the first parameter combination is used to characterize model parameters of the initial model in compliance with the constraint model; Update the initial model according to the first parameter combination to obtain a first model; Based on the first model, determining a second parameter combination using the target algorithm, wherein the second parameter combination is a guiding parameter determined by the target algorithm and is used to adjust the model parameters of the first model; The target model is determined according to the second parameter combination and the first model.
5. The base station load balancing method according to claim 4, characterized in that: The target training comprises the following steps: Adjusting the initial parameters of the initial model based on the i-th historical service combination to obtain a third parameter combination; Update the initial model according to the third parameter combination to obtain a second model; Training the i-th historical service combination according to the second model to generate a first strategy; Determine whether the third parameter combination is the first parameter combination according to the first strategy and the constraint model.
6. The base station load balancing method according to claim 5, characterized in that: Determining whether the third parameter combination is the first parameter combination according to the first strategy and the constraint model includes: Processing the i-th historical service combination according to the first strategy to obtain a processing result, wherein the processing result at least includes a performance indicator for completing the i-th historical service combination; Determining whether the processing result conforms to the constraint model; When the processing result meets the constraint model, it is determined that the third parameter combination is the first parameter combination of the initial model.
7. The base station load balancing method according to claim 1, characterized in that: After adjusting the target model based on the target service combination and generating a load balancing strategy for the target service combination based on the adjusted target model, the method further includes: Allocating resources to the target service combination according to the load balancing strategy to obtain an allocation result; Executing the target service combination based on the allocation result to obtain an execution result, wherein the execution result at least includes a performance indicator for completing the target service combination; The target model is updated according to the execution result.
8. A base station load balancing device, characterized in that: include: A prediction unit, used to predict the distribution characteristics of network service demand of a target scenario within a target time period based on historical service data; An adjusting unit, configured to adjust the resource distribution of the base station according to the distribution characteristic; A receiving unit, configured to receive a target service combination within the target time period, wherein the target service combination includes M services and a priority of each of the services, wherein M is an integer greater than or equal to 1; A generating unit is used to adjust the target model according to the target service combination and the target algorithm based on the resource distribution of the base station, and generate a load balancing strategy for the target service combination based on the adjusted target model, wherein the target algorithm is used to adjust the model parameters of the target model based on the target service combination without adjusting the structure of the target model, so that the target model generates a load balancing strategy that conforms to a constraint model, wherein the constraint model at least includes a policy target, wherein the policy target is used to characterize that a service rate of a service with a priority less than a preset value among the M services is greater than a preset rate, and that a service time variance of the target service combination is less than a preset threshold.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the base station load balancing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the base station load balancing method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Hierarchic scheduling method for service and resources in cloud computation environment
CN103701886A
Intelligent power grid load prediction method based on federated learning
CN116706888A
Cloud edge computing task scheduling method based on reinforcement learning
CN118740835A
Load balancing optimization method and device, electronic equipment and readable storage medium
CN118939414A
Scheduling automation system application state management method
CN119292745A