Method and apparatus for ran optimization using heuristic-initiated reinforcement learning
The heuristic-initiated reinforcement learning method addresses the challenges of large-scale RAN optimization by using heuristic recommendations to train AI/ML models, ensuring reliable and explainable optimization, enabling immediate deployment and continuous training while minimizing network degradation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MAVENIR SYST INC
- Filing Date
- 2025-10-14
- Publication Date
- 2026-04-23
AI Technical Summary
Current RAN optimization methods, particularly in 4G and 5G networks, face challenges in large-scale live deployments due to the impracticality of manual configuration, lengthy training cycles of reinforcement learning (RL) models, unpredictable outcomes, and the need for reliable and explainable optimization techniques.
A heuristic-initiated reinforcement learning method that parallelly operates with AI/ML RL, using heuristic recommendations to train AI/ML models, ensuring convergence and providing a failsafe mechanism to switch back to heuristics during issues, allowing for gradual exploration of unexplored configurations.
Enables immediate deployment with reliable and explainable optimization, reduces the risk of network degradation, and allows for continuous training and adaptation to changing conditions, ensuring high-quality service and efficient resource utilization.
Smart Images

Figure US2025050775_23042026_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR RAN OPTIMIZATION USING HEURISTIC- INITIATED REINFORCEMENT LEARNINGBACKGROUND OF THE DISCLOSURE1. Field of the Disclosure
[0001] The present disclosure is directed to radio access network (RAN) optimization, particularly for 4G and 5G based mobile networks, and more particularly, optimization for such networks in live deployments at scale.2. Description of Related Art
[0002] The rapid expansion of mobile communication networks, particularly with the deployment of 4G and 5G technologies, has created a need for optimizing RANs.
[0003] RAN optimization refers to the processes and activities required to improve or maintain the performance of an already active or operational wireless communication network. These optimizations function to ensure high-quality service and efficient utilization of network resources in real-time environments. However, in live network deployments, operators face significant challenges when experimenting with new optimization techniques.
[0004] For example, in a telecommunications environment, careful configuration of numerous parameters is essential to avoid unintended impacts on network performance or unexpected degradation of the user experience. Any failure in a live network can result in significant consequences for commercial operations. As a result, telecom operators rigorously evaluate proposed changes before implementing any adjustments to network configurations or parameters.
[0005] Manually optimizing RANs is impractical for large-scale networks with RAN deployments of thousands of sites, and it becomes infeasible to manually manage more than a few sites individually.
[0006] To optimize RAN deployments at scale and in an automated manner, SelfOrganizing Networks (SON) have been traditionally employed. SONs are RANs that are designed to autonomously plan, configure, manage, optimize, and heal themselves. SONs offer automated functions such as self-configuration, self-optimization, self- healing, and self-protection.
[0007] Self-configuration aims to achieve a "plug-and-play" model, where new base stations are automatically configured and integrated into the network. This includes both establishing connectivity and downloading configuration parameters via software.Typically, equipment vendors provide self-configuration as part of the software package for each radio cell. When a new base station is introduced and powered on, it is instantly recognized and registered by the network. Neighboring base stations then automatically adjust their technical parameters, such as transmission power and antenna tilt, to ensure optimal coverage and capacity while minimizing interference.
[0008] Self-optimization aims to autonomously and continuously adapt system settings to a traffic profile and the network environment in terms of topology, propagation, and interference. Each base station is equipped with hundreds of configuration parameters that control various aspects of the cell site’s operation. These parameters can be adjusted to modify network behavior based on observations from both the base station and measurements collected from mobile devices.
[0009] Self-healing aims to mitigate an impact by adjusting parameters and algorithms in adjacent cells, allowing other nodes to take over the load from a failed node. In legacy networks, identifying malfunctioning base stations can be challenging, often requiring considerable time and resources to address the issue. With SONs, this function enables the immediate detection of failing base stations, allowing for prompt corrective measures and minimizing or even preventing service degradation for users.
[0010] Self-protection aims to ensure the system's security remains impenetrable and to safeguard the confidentiality and integrity of the data, for example, to defend against unauthorized access and protect against both active and passive attacks.
[0011] These systems typically rely on heuristic-based methods, where rules or algorithms are created and pre-defined based on specific network conditions to initiate and provide configuration changes that improve network performance.
[0012] With advancements in artificial intelligence and machine learning (AI / ML), there is a growing interest in leveraging these technologies to enhance RAN optimization further. Techniques such as reinforcement learning (RL) are gaining traction, allowing network entities to learn from interactions within the network environment, adapt to changing conditions, and improve performance based on the outcomes of previous actions.
[0013] Current RL methods for model training typically require a long training cycle (e.g., several months of field training) and validation for every optimization use case. Such a lengthy process is not practical in many scenarios, such as during an expansion phase of a live network deployment, where rapid optimization is often desirable.
[0014] Field training using RL involves a complete and extensive exploration of the parameter search space, which can result in suboptimal choices and lead to key performance indicator (KPI) degradation during the training process. In a live deployment, such a method may not be feasible, as the choices could negatively affect user experience during the training period. The expectation from operators is to have outcomes that can be explained and understand the benefit of the change before it can be applied in the network.
[0015] It is also often necessary to establish or determine a baseline performance of the AI / ML models since it is not possible to reliably predict whether KPI improvements achieved by using RL can outperform traditional optimization methods.
[0016] There may also be times when the input data used to train the AI / ML models becomes corrupted or when retraining is required due to changes in site deployments, configurations, or traffic patterns caused by factors like new devices or special events. In these situations, fallback mechanisms are needed to ensure continuous network performance.
[0017] Accordingly, it has been determined by the present disclosure that there is a continuing need for RAN optimization methods and apparatuses that overcome, alleviate, and / or mitigate one or more of the aforementioned and other deleterious effects of the prior art.SUMMARY OF THE DISCLOSURE
[0018] The present disclosure provides a method for supporting the parallel operation of a heuristics engine and AI / ML RL to facilitate the deployment of SON using RL in live commercial networks.
[0019] The present disclosure has found that, while offline reinforcement learning techniques can be used to initiate the learning process, they may not fully meet the reliability and confidence requirements of telecom operators in a live deployment, and that a cautious method leveraging heuristics is desirable. The outcomes of heuristicbased methods are explainable and can be reasoned through RAN domain expertise. By using heuristics, random or unpredictable outcomes are less likely, allowing operators to build confidence in their use based on clear reasoning without relying on machine-driven decisions until AI / ML models achieve the necessary confidence levels.
[0020] The present disclosure provides a heuristic-initiated RL method for radio network optimization in live deployments, where the inputs to the RL method are simultaneously fed to the heuristic engine as well. Initially, recommendations generated by using heuristics are evaluated and used to train AI / ML RL models. Then, during inference, the performance of the AI / ML model is periodically checked at set time intervals, and the outputs from the RL are compared with the outputs generated by the heuristic engine. Once the AI / ML models converge to meet or exceed the performance of the heuristic engine, for example, based on a predefined convergence criterion, the AI / ML models can operate in place of the heuristic engine. Such a method allows for additional training phases of the AI / ML models that can explore the solution space for further improvements. Further improvements can be performed in a more cautious manner by focusing on the good or more improved configurations determined by theheuristic engine. This is possible since the baseline model is trained to at least match the heuristics, which provides a foundation for further improvements.
[0021] This method allows the system to be confidently deployed in the field even before the training is fully completed, while simultaneously continuing to train the AI / ML RL models. The AI / ML RL models can be trained in parallel to the operation of the heuristics engine. Additionally, the method provides a failsafe in the field, ensuring that in the event of AI / ML performance issues, product malfunctions, or network problems, the system can switch to heuristic-generated configurations for a certain period of time as necessary. The heuristics engine can also be employed during retraining phases, such as those required for site expansions or when model failures occur due to bad, invalid, or incorrect data or a change in environment conditions.
[0022] The present disclosure provides a method in which heuristics can also be used to guide the behavioral policy for offline RL. Given the high cost and risk associated with direct interactions in real-world environments, offline RL, which learns from pre-collected datasets, can be a more practical method for real-world deployments.
[0023] The present disclosure provides a heuristic-initiated reinforcement learning method for radio network optimization, where the behavioral policy is dictated by the heuristics. Recommendations from the heuristics engine are evaluated and applied to the network functions (NFs), and the resulting feedback data is used to train AI / ML RL models. Periodically, the performance of the AI / ML model’s output is evaluated during inference and compared with the heuristics.
[0024] In addition to relying on the behavioral policy used by the heuristics engine for implementing configuration changes and generating observability feedback data while exploring the action search space, the method also employs RL-based exploration and its distinct behavioral policy, different from the behavioral policy employed by the heuristic engine. This method allows for additional configuration changes that delve into previously unexplored action space values by the heuristics engine to gather further feedback data.
[0025] To balance the need to explore action parameter values that are not explored by the heuristics engine while minimizing the unpredictability of RL-based exploration, the method uses a weighted method. A relative weight ratio is applied between the heuristic policy and the RL-based exploration. This allows alternating between the heuristic-based and RL-driven behavioral policies when making configuration changes on the network functions (NF) and gathering feedback data. The present disclosure has found that such a method enhances diversity and heterogeneity across feedback observability data from the NFs by incorporating multiple behavioral policies, providing richer data for network optimization.
[0026] Once the AI / ML models converge on the target policy based on a specific criterion, such as Bellman Loss, Temporal difference (TD) error, or Log-likelihood loss, during training, these trained models can replace the heuristics-based and RL-based exploration. This allows for further bootstrapping of additional training phases, occasionally exploring the solution space for further improvements. The AI / ML models leverage the learning from the target policy to generate more informed inferences, which are expected to be at least as effective as the heuristics-based output while also identifying configuration regions that may yield beneficial results and have not yet been explored by the heuristic engine.
[0027] This method provides a failsafe mechanism of using heuristics instead of relying on random exploration in combination with an exploration of the search space, controlled by a configurable trade-off weight. This can be done by starting with a known value, which is conservative but may not be optimal. The method can then move in steps, for example, 10% in either direction until an acceptable maximum value is achieved. This weight is configurable and reflects the deviation from the known value. Additionally, this method can be employed during retraining phases, which may be desirable during site expansions and to address model failures due to bad, invalid, or incorrect data.
[0028] A method for optimizing a RAN according to the present disclosure includes the steps of: a. generating, by a heuristic server unit, a heuristic recommendation foroptimizing network configurations with a heuristic-based algorithm using data from the RAN, such as counters, performance measurements, traces, and KPIs as input; b. applying, by an Operations, Administration, and Maintenance (OAM) unit, the heuristic recommendation as a configuration to the RAN; c. collecting, by an AI / ML server unit, feedback data from the RAN after applying the heuristic recommendation; d. using, by the AI / ML server unit, the collected feedback data to train an AI / ML RL-based model; e. generating, by the AI / ML server unit, an RL recommendation for optimizing network configurations of the RAN with the AI / ML RL-based model using data from the RAN as input; f. comparing, by a comparison unit, the heuristic recommendation to the RL recommendation to determine whether a convergence is achieved; and g. one of: (i) switching, by a switching unit, to RL recommendations if convergence is achieved and applying future RL recommendations as a configuration to the RAN, or (ii) if convergence has not been achieved in step f, repeating steps a to f until convergence is achieved.
[0029] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the method also includes alternatingly repeating both steps a and b, and then step e.
[0030] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the method also includes performing both steps a and b and then step e at a ratio based on a behavioral policy determined by the heuristic-based algorithm of the heuristic server unit to make changes to the network functions of the RAN using the OAM unit and collecting feedback data from the RAN to provide as input to the heuristic server unit and the AI / ML server unit.
[0031] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, convergence is achieved when the most recent heuristic recommendations are substantially the same as previously generated heuristic recommendations.
[0032] In some embodiments according to the disclosure, either alone or togetherwith any one or more of the aforementioned and / or after-mentioned embodiments, convergence is achieved when at least one of the following conditions is achieved: (i) when the most recent heuristic recommendations are substantially the same as previously generated heuristic recommendations, and (ii) the AI / ML is generating recommendations that are within a predetermined percentage of actions recommended by heuristics within a predetermined time period.
[0033] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the method is performed on a live RAN deployment with the heuristic server unit and the AI / ML server unit operating simultaneously in parallel.
[0034] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the heuristic-based algorithm determines action states with good outcomes that can be used to generate the heuristic recommendation.
[0035] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the method also includes training the AI / ML RL-based model using all the action states determined to have good outcomes.
[0036] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the method also includes determining action states with bad outcomes and excluding those action states from being applied in deployment to train the AI / ML RL-based model.
[0037] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, step e further comprises performing a neighbor search around good outcomes.
[0038] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, step e further comprises transfer learning.
[0039] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, step e is performed as a simulation in a lab environment.
[0040] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the method also includes performing step h until a trigger occurs, and then repeating the method.
[0041] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the method also includes the trigger is at least one selected from the group consisting of: system parameters exceeding predefined thresholds indicating suboptimal performance, a change in data set, a change in environment conditions, a new software deployment that changes traffic patterns, a new device introduced in the network, an addition of a new cell site, and a change in channel geometry due to new buildings
[0042] A system for optimizing a radio access network includes a heuristic engine configured to generate behavioral policy recommendations for radio network configuration, an AI / ML reinforcement learning engine trained using feedback data collected from the application of heuristic-based recommendations on network functions of the RAN and configured to generate behavioral policy recommendations for radio network configuration, a model inferencing engine that compares an output of the reinforcement learning model with heuristic-based recommendations during real-time operations, a switch configured to switch between heuristic-based policies and outcomes from the heuristic engine and RL-based exploration from the AI / ML reinforcement learning engine, and a management system configured to apply the behavioral policy recommendations as configuration for network functions of the radio access network.
[0043] A method for dynamically switching between an RL-based model and a heuristic-based algorithm to optimize a radio access network (RAN) according to the present disclosure includes the steps of: (a.) generating, by an AI / ML server unit, an RLrecommendation for optimizing network configurations of the RAN with an AI / ML RL- based model using data from the RAN as input to the model; (b.) monitoring, by a comparison unit, one or more system parameters related to performance of the AI / ML RL-based model to determine whether a trigger condition has occurred; and (c.) one of (i) switching, by a switching unit, to generating, by a heuristic server unit, a heuristic recommendation for optimizing network configurations with a heuristic-based algorithm and applying by an Operations, Administration, and Maintenance (OAM) unit, the heuristic recommendation as a configuration to the RAN if the trigger is determined to have occurred, or (ii) applying, by the OAM unit, an RL recommendation as a configuration to the RAN if the trigger is determined to not have occurred.
[0044] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the trigger is at least one selected from the group consisting of: system parameters exceeding predefined thresholds indicating suboptimal performance, a change in data set, a change in environment conditions, a new software deployment that changes traffic patterns, a new device introduced in the network, an addition of a new cell site, and a change in channel geometry due to new buildings.
[0045] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the method also includes, after switching to generating the heuristic recommendation in step c, monitoring for a presence of the trigger condition, and if no longer present, switching, by the switching unit, back to generating RL recommendations by the AI / ML server unit.
[0046] In some embodiments according to the disclosure, either alone or together with any one or more of the aforementioned and / or after-mentioned embodiments, the method also includes defining, by the switching unit, an upper threshold and a lower threshold for the trigger and only switching between the RL-based model and the heuristic-based algorithm when the upper threshold or lower threshold is exceeded.
[0047] In some embodiments according to the disclosure, either alone or togetherwith any one or more of the aforementioned and / or after-mentioned embodiments, the switching in step c occurs only if the trigger is determined to have occurred and persists for a predetermined period of time.
[0048] The above-described and other features and advantages of the present disclosure will be appreciated and understood by those skilled in the art from the following detailed description, drawings, and appended claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0049] FIG. 1 A shows an example D-SON.
[0050] FIG. 1 B shows an example C-SON.
[0051] FIG. 1 C shows an example AI / ML-based method for RAN optimization.
[0052] FIG. 2 shows a representative example of RL outcomes.
[0053] FIG. 3 shows a representative example RL model training with two phases.
[0054] FIG. 4 shows a representative example method to prune the action space forRL with only known good action outcomes.
[0055] FIG. 5 shows a representative example method to mitigate the risk of bad outcomes for RL.
[0056] FIG. 6 shows a diagram of an example heuristic-initiated RL method according to the present disclosure.
[0057] FIG. 7 shows the expected benefits of the heuristic-initiated RL method compared to a baseline SON implementation in terms of performance vs. time.
[0058] FIG. 8 shows a representative deployment of a Radio Access Network (RAN) system that integrates SON optimization capabilities.DETAILED DESCRIPTION OF THE DISCLOSURE
[0059] The present disclosure provides improvements over various methods for RAN optimization using SONs.
[0060] A plurality of Distributed SONs (D-SON) 105 is shown in FIG. 1A. D-SONs 105 relay commands across the network edge, where each node exchanges information with the others. The D-SONs are located at each eNodeB (eNB) or gNodeB (gNB) and have knowledge of both cell and user equipment (UE) information within the cell, allowing for rapid response. High-level reports 101 are sent to the management system, Operations, Administration, and Maintenance (0AM) 104, which can configure certain policies 102 to influence D-SON optimization settings.
[0061] This decentralized architecture offers greater flexibility compared to centralized SON and enables faster response times and visibility of network changes. With this architecture, individual node optimization does not always guarantee overall network improvement, as each node operates independently as a relay.
[0062] A Centralized SON (C-SON) is shown in FIG. 1 B. Functions are managed at higher-level network nodes. In this architecture, performance counters and KPIs 123 are sent to the C-SON node, and the resulting configurations 124, such as commands, adjustments, settings, and requests, are distributed directly from the network management console and then propagated to each individual node. This centralized method allows for a comprehensive view of the entire network by scanning and taking in global aspects of the network, enabling the algorithm to make more informed decisions about optimizations and future configurations. This setup has slightly longer response times compared to a D-SON, where a speed trade-off provides increased control and precision in managing network resources.
[0063] An AI / ML SON implementation is shown in FIG. 1 C. AI / ML technologies enable SON to become more intelligent, adaptive, and scalable by allowing the network to learn from data, such as measurements, statistics, and feedback. Al and ML can apply advanced algorithms, including optimization, classification, and prediction, toautomate and assist in the decision-making process. The availability and quality of data are critical for training and validating AI / ML models. The data must be representative, reliable, consistent, and compliant with privacy and security regulations. Additionally, the complexity and scalability of AI / ML algorithms must handle the large-scale, multidimensional, and non-linear nature of network problems while delivering robust and interpretable results.
[0064] RL is a type of machine learning in which a software agent learns how to behave in an environment to achieve a specific goal, such as maximizing a reward. This makes RL particularly well-suited for control implementations like those found in RAN optimization. RL operates with three core components: state, action, and reward. The state, “s”, is a set of values, or features, that characterize the environment in a way that is relevant to the problem. The action “a” represents the change the agent applies to the environment, while the reward “r” is a multi-objective scalar function that numerically expresses the agent's objective
[0065] The agent’s interaction with the environment is captured over time through samples of experiences represented as tuples (st, at, rt+i , St+i ), where “t” is a discrete time step, describing a state transition as a consequence of applying actions to the environment and receiving rewards. The object of RL is to learn and extract from these samples a policy “p” that, given a state, returns the optimal action to be taken to maximize the long-term cumulative reward. An RL algorithm maps past rewards to previous actions taken, rapidly guiding the agent from an initial state of uncertainty, where it does not know how to act, to a state of near-optimal performance where it acts as close to optimality as possible.
[0066] By way of a non-limiting example, to optimize energy usage in a dense urban 5G network while ensuring good user experience, e.g., high throughput, the following can be applied. The state includes the current traffic load on a specific base station, the signal strength for different users, the energy consumption of the base station, and the interference levels. The action can be to turn off some base station antennas when the traffic load is low to save energy. The reward can be a combination of factors, such asreducing energy consumption, i.e., a positive reward, and ensuring that user throughput stays above a certain threshold. The reward becomes negative if it drops below a certain threshold. The policy could then be a dynamic power control strategy that reduces power during off-peak times while still maintaining adequate coverage and quality for users. Thus, a 5G network could dynamically decide which cells to shut down based on varying conditions and optimize for performance, efficiency, and user satisfaction.
[0067] Training is the first phase of developing an Al model, during which the model learns from labeled data to generate task-specific parameters, such as weights and biases. This process involves presenting the model with input examples and their corresponding outputs, allowing it to adjust its internal parameters using optimization techniques like gradient descent to minimize prediction errors. Once the training is complete, the model is capable of making predictions or classifications based on new, previously unseen data.
[0068] Inference is the second phase, which occurs after training. Inference refers to the process of using the pre-trained model to make predictions or decisions based on input. During this phase, the trained model is put into practical use for real-time decision-making or predictions, whereby the model processes new input data and produces an output or conclusion without human intervention.
[0069] FIG. 2 shows a representation of an RL 200. RL 200 involves exploring action states by an agent to discover how the different action states affect the RAN and what rewards might produce and can yield either good configurations 201 or bad configurations 202. Good configurations 201 are good results from exploration, whereas bad configurations 202 are bad results from exploration. Good configurations 201 are more optimal configurations, settings, or policies that were not previously known or used, whereas bad configurations 202 are less optimal than previously known or used.
[0070] In an example embodiment of the present disclosure, heuristics are used to explore the good configurations and train the AI / ML model while the system is initially deployed and actively used in the field. Unknown configurations are explored using avariety of techniques, including but not limited to a neighborhood search of the good configurations from heuristics, transfer learning, simulations, and / or lab testing. The unknown configurations are better, but also potentially bad configurations that are then determined to be either good configurations 201 that can be exploited or bad configurations 202.
[0071] In another example embodiment according to the present disclosure, shown in FIG. 3, the RL model training process 300 has two phases. The first phase is a heuristics phase 301 , which is SON-based. The second phase is an AI / ML phase 302, which is RL-based. In this example, training starts in heuristics phase 301. After convergence 303, the training switches to AI / ML phase 302.
[0072] During AI / ML phase 302, if an issue with the AI / ML model is identified in which retraining is required or a trigger 304 occurs, the training switches back to heuristics phase 301 , and the process repeats, i.e., training restarts in heuristics phase 301 until subsequent convergence 303.
[0073] Examples of issues that, if identified, result in or trigger falling back to retraining in heuristics phase 301 include but are not limited to changes in data sets, changes in network configuration, changes in traffic patterns, and changes in the number of sites.
[0074] According to another example embodiment of the present disclosure shown in FIG. 4, a method is provided that uses the heuristics to prune the action space for RL with only known good action outcomes. This is critical because, in real-world field deployments, a bad choice of outcome can significantly degrade the user experience, leading to unacceptable performance issues such as dropped calls, slower data speeds, or poor connectivity. These types of failures are highly undesirable for network operators, who prioritize maintaining high levels of service reliability and user satisfaction. A single bad decision in configuration or optimization could result in widespread disruptions, impacting a large number of users and potentially leading to customer dissatisfaction, service complaints, or even financial losses.
[0075] In the heuristics phase shown in FIG. 4, data and configuration information from the RAN are fed to both the SON heuristics engine and the AI / ML-based RL model, thus enabling the training of the model in parallel to using the heuristics engine.
[0076] Step A shows the action space to be explored for RL. When the data is fed to both the heuristics engine and the AI / ML training, the heuristic provides a first configuration recommendation 401 while the RL provides a second recommendation 402, as shown in step B. Second recommendation 402 can be identical to first configuration recommendation 401 or different from first configuration recommendation 401.
[0077] The SON heuristics engine can obtain a good outcome, for example, first recommendation 401 , which, even if not necessarily optimal, can still be good enough for implementation in the field. In such a case, second recommendation 402 from the RL can be discarded. First recommendation 401 can be applied to the RAN.
[0078] Even though second recommendation 402 is discarded, the RL model receives first configuration recommendation 401 and data based on the application of first configuration recommendation 401 to the RAN as input. Accordingly, the RL model is aware of the network configuration change and first configuration recommendation 401 of the heuristics engine, along with its resultant good performance, to continue training and further explore the action space.
[0079] The process continues as shown in step C, where the SON heuristics engine is used to continue to explore the action space until a convergence criterion is satisfied.
[0080] Upon completion of a further convergence, the RL will be trained on good outcomes while some remaining action states remain unexplored. As shown in step D, these unexplored action states may have better outcomes 403 than what has been explored. Conversely, the unexplored action states may also have some very bad outcomes 404. Predictions on the type of outcome can be made by the AI / ML.
[0081] In FIG. 5, various options are shown that can be explored to mitigate the risk of bad or very bad outcome states while exploring further states by the AI / ML.
[0082] Step A shows the action states that were identified during heuristics phase 301 as good outcomes 501 , while the remaining action states remain unexplored. Step B shows the remaining unexplored states, while good outcomes 501 have been pruned from the RL exploration. To further prune the RL and explore good states as shown in step C, various pruning options can be explored so that at the end of the exploration, there remain unexplored states 502 that are not to be explored in the field or online due to the high risk of poor outcomes, and thus undesirable consequences of degraded network performance.
[0083] A first pruning option is a neighbor search around identified good outcomes. In this method, the unexplored action states are searched based on a neighborhood of the heuristics provided, with the expectation that the action states close to good outcomes determined by SON heuristics are likely to also be good. This can be used to further refine the search space, and the range of search can be gradually increased depending on the outcome. For example, neighboring cells or base stations (e.g., eNodeBs in 4G or gNodeBs in 5G) that are in close proximity to a particular cell can be identified.
[0084] A second pruning option is transfer learning. In this method, training is based on changes in the input data that are then correlated to known action states without actually having to configure the recommendations in the field to validate the outcomes.
[0085] Transfer learning allows models to leverage previously learned patterns, parameters, or policies, significantly reducing the time and data needed for new training phases. The model’s knowledge is transferred from a source domain where the model has been previously trained and can be adapted to a target domain where the model will be applied.
[0086] For example, a trained model from a source domain could have learned optimization strategies for a 4G network with specific traffic patterns, geographic conditions, or network topology, and could be applied to a target domain that is a 5G network, a new geographic area, or has different traffic patterns. As another example, the source domain strategy can be applied to a target domain that has different, newerdevice profiles and newer applications that present different traffic and signaling characteristics.
[0087] A third pruning option is simulations. In this method, unexplored states can be simulated in a lab environment, offline. The outcomes of the simulations can then be used to train the RL model.
[0088] Figure 6 shows an example method 600 according to the present disclosure of a parallel operation of both a heuristic and RL training with switching conditions between both the heuristic-based and the AI / ML RL-based methods.
[0089] RAN NF 610 generates counters, performance measurements, traces, and KPIs, which are fed to RL engine 601 .
[0090] RL engine 601 includes a heuristics engine 620 and AI / ML engine 630 that operate in parallel. An output of RL engine 601 switches between an output of heuristics engine 620 and an output of AI / ML engine 630.
[0091] Heuristics engine 620 can be, for example, a heuristic server unit comprised of logic and circuitry. A heuristic server unit can have an input interface, including a data acquisition module that gathers input from various sources such as databases, sensors, or external systems. This module pre-processes and filters the data to make it suitable for heuristic processing. The heuristic server unit can also have a client input interface that allows other systems to interact with the server, for example, via an Application Programming Interface (API), a Graphical User Interface (GUI), or a Command Line Interface (CLI). The heuristic server unit can have one or more Central Processing Units (CPU) that execute heuristic algorithms, including rules-based logic and heuristic searches. The heuristic server unit can have non-transitory storage for heuristic algorithms, rules, or models the system uses to make decisions or solve problems, and a database that stores pre-computed heuristics, patterns, or historical data for the system to reference and guide its decision-making process. The heuristic server unit can have logic to apply the stored heuristic rules and algorithms to input data, processing it to produce decisions or recommendations, and / or a pattern recognition0016958WQU / 4688 engine to detect patterns in the input and apply the most appropriate heuristic rule or action. The heuristic server unit can also have an output interface whereby processed output is passed through the interface to provide results to downstream systems, in the form of recommendations, decisions, or optimized solutions. The heuristic server unit can also have a network interface to manage communication between the heuristic server unit and other systems or servers.
[0092] The components of the heuristics server unit can either be centralized in a single system or distributed across multiple systems. The advantage of a distributed system is that the components of the heuristics server unit are separated across multiple physical or virtual servers, often located in different regions or data centers. These servers communicate over a network and can scale horizontally as needed.
[0093] AI / ML engine 630 can be, for example, an AI / ML server unit, which is a specialized computing system designed to train, deploy, and serve artificial intelligence (Al) and machine learning (ML) models. AI / ML server unit can handle tasks such as data preprocessing, model training, inference, and monitoring. The AI / ML server unit can include a data acquisition module to collect raw data from various sources, a data processing unit to clean and transform raw data into a suitable format for model training or inference, and a data validation module that ensures the integrity and quality of the data before it's used in modeling. The AI / ML server unit can include data storage with non-transitory memory, including databases or data lakes, e.g., SQL databases, NoSQL databases, Hadoop distributed file system (HDFS), to store raw and processed data. The AI / ML server unit can also have a model storage repository that stores trained models along with their versions, metadata, and performance measurements for easy retrieval and deployment. The AI / ML server unit also includes various processing units, including CPUs or general-purpose processors for handling standard computations, GPUs, preferably optimized for model training, Tensor Processing Units (TPU) that are specialized hardware accelerators, and Field-Programmable Gate Arrays (FPGA). The AI / ML server unit can have a model training engine that orchestrates the training process, managing data pipelines, model configurations, and hyperparameter tuning. The AI / ML server unit can have an experimentation platform that facilitatesexperimentation with different models and parameters, for example, using weights and biases. The AI / ML server unit can include various machine learning frameworks and libraries. The AI / ML server unit can include an inference engine to make predictions on new data, a model serving engine to deploy models, and a model modeling engine to observe model performance metrics like accuracy, latency, and drift detection to ensure models remain effective over time. The AI / ML server unit can include a network interface to manage communication between the AI / ML server unit and other systems or servers.
[0094] Various deployments of the AI / ML server are possible. Non-limiting examples include a centralized deployment in which all components reside within a single server or data center, a distributed deployment in which components are spread across multiple servers or server clusters, a cloud-based deployment, and a hybrid deployment that combines on-premise infrastructure with cloud resources.
[0095] Counters are raw data points or statistics collected by network equipment (such as base stations, eNodeBs, or gNodeBs) over time. They represent specific events or occurrences within the network, such as the number of successful or failed attempts for particular operations that are generated or pegged by the applications. Counters provide granular, real-time data about various aspects of network performance. For example, counters can count how many times a mobile device successfully transitions between base stations, track how many active connections were dropped unexpectedly, measure how many calls were initiated in a specific cell or area, measure how many data packets were lost or corrupted during transmission, track CPU resources being utilized by a base station or network node, or indicate how many times users attempted to connect to the network.
[0096] Traces encompass signaling events throughout the radio network interfaces between the cell phone and the corresponding network elements, providing critical Radio Frequency (RF) parameters such as cell signal strength (Reference Signal Received Power or RSRP) and quality (Reference Signal Received Quality or RSRQ).
[0097] Performance measurements (PMs) translate the counters into meaningfulmeasurements, typically in a standardized manner that can be used to help indicate the performance of the network.
[0098] KPIs are high-level metrics derived from performance measurements and are used to evaluate the overall performance and quality of service (QoS) of a network.They provide a summary of how well the network is functioning in relation to specific goals or benchmarks, such as user experience, network quality, and service availability. Network operators typically monitor KPIs to ensure the network is meeting performance targets. Examples of KPIs include call setup success rate (CSSR), handover success rate (HOSR), dropped call rate (DCR), throughput, latency, interference levels, cell availability, resource block utilization, peak hour traffic, data volume, jitter, voice quality, video streaming quality, energy consumption per bit, network congestion, RACH Success Rate, PDU session establishment and modification success rate, RSRP / RSRQ distribution, Radio Link Failures, and the like.
[0099] Both the heuristics engine 620 and AI / ML engine 630 phases are operated in parallel. Selection of the output of RL engine 601 vs. heuristics engine 620 is determined by a switch 640 in conjunction with KPI monitoring engine 650. Switch 640 is a switching unit.
[0100] Initially, switch 640 is set to output the recommendation from heuristics engine 620 to 0AM management system 660, and RL is performed by AI / ML engine 630, learning about the good action outcomes based on what the heuristic algorithm of heuristics engine 620 is proposing.
[0101] 0AM management system 660 applies the recommendations output by switch 640 as configurations to RAN NF 610.
[0102] KPI monitoring engine 650 is a comparison unit that actively monitors the results of heuristics engine 620 and AI / ML engine 630 to determine which results to output via switch 640. Upon convergence of heuristics engine 620, output switches to AI / ML engine 630. Upon identification of an issue or occurrence of a trigger, the output switches back to heuristics engine 620.
[0103] Convergence of the heuristics engine 620 is achieved when one or more of the conditions below are satisfied:1 . Heuristic action output recommendations from heuristics engine 620 are not changing significantly. By way of a non-limiting example, the recommendation values are converging and stay within 5% of previous recommendations;2. The RL algorithm of AI / ML engine 630 is selecting actions that are similar to those that the heuristic algorithm of heuristics engine 620 is selecting within certain bounds. By way of a non-limiting example, the bounds can be 5% for 80% of the selections within a time period of 24 hours, all of which are configurable; and3. A manual switch is performed based on known maintenance or system configuration changes.
[0104] Upon convergence, the switch is made to the RL algorithm of AI / ML engine 630. Upon the occurrence of certain triggers, the RL algorithm of AI / ML 630 switches back to heuristics engine 620.
[0105] In example embodiments, the triggers should persist for some configurable duration of time.
[0106] Non-limiting example triggers include changes in data sets, performance outside expected bounds, changes in environment conditions, new software deployments that change traffic patterns, new devices introduced in the network, additions of new cell sites, or changes in channel geometry due to new buildings.
[0107] In example embodiments, a hysteresis threshold is applied to avoid the ping- pong effect and prevent frequent or unnecessary switching between heuristics engine 620 and AI / ML engine 630.
[0108] FIG. 7 shows the expected benefits of the heuristic-initiated RL method compared to a baseline SON implementation in terms of performance vs. time. It enables the deployments to start immediately using heuristics without waiting fortraining of the AI / ML schemes to be completed. The AI / ML training continues in parallel to the heuristic-based method. The goal is first to select the “good” configurations during deployment and then cautiously explore the unknown configurations to achieve performance comparable to traditional RL exploration.
[0109] FIG. 8 shows a representative deployment of a Radio Access Network (RAN) system that integrates SON optimization capabilities. In this configuration, the deployment at a typical cell site 801 includes antennas, which are responsible for transmitting and receiving wireless signals, radio units (RU) that modulate and demodulate those signals, and the baseband processing that takes place at the virtualized Distributed Unit (vDU). The vDU handles tasks such as signal encoding, decoding, and other real-time processing functions to ensure seamless communication within the network.
[0110] Further upstream in the architecture, a virtualized Centralized Unit (vCU) is responsible for managing higher-layer protocol functions. The vCU is typically hosted within a local data center 802. An Element Management System (EMS) provides network management and monitoring functionalities, facilitating the operational oversight of the deployed infrastructure. The EMS is also hosted within local data center 802.
[0111] The core network, which is responsible for routing data between the RAN and external networks, is typically situated in a larger regional or national data center 803, where more substantial computing resources are available. This separation of components allows for enhanced scalability, efficiency, and flexibility in managing network operations.
[0112] SON optimization 804 provides automation of network operations, utilizing the combination of heuristic methods and AI / ML algorithms according to the present disclosure to optimize the various network functions. SON optimization 804 can be hosted, for example, in a centralized location like local data center 802 to maximize the ability to collect and analyze data from various sites. One or more Graphics Processing Units (GPU) 805 are used for the offline training of the AI / ML models according to thepresent disclosure.
[0113] It should also be noted that the terms “first”, “second”, “third”, “upper”, “lower”, and the like may be used herein to modify various elements. These modifiers do not imply a spatial, sequential, or hierarchical order to the modified elements unless specifically stated.
[0114] As used herein, the term “substantially” refers to the complete or nearly complete extent or degree of an action, characteristic, property, state, structure, item, or result. The exact allowable degree of deviation from absolute completeness can, in some cases, depend on the specific context. However, generally speaking, the nearness of completion will be to have the same overall result as if absolute and total completion were obtained.
[0115] As used herein, the term “about” is used to provide flexibility to a numerical range endpoint by providing that a given value may be “a little above” or “a little below” the endpoint. Further, where a numerical range is provided, the range is intended to include any and all numbers within the numerical range, including the endpoints of the range.
[0116] While the present disclosure has been described with reference to one or more examples, it will be understood by those skilled in the art that various changes can be made, and equivalents can be substituted for elements thereof without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt a particular situation or material to the teachings of the present disclosure without departing from the scope thereof. Therefore, it is intended that the present disclosure will not be limited to the particular examples disclosed as the best mode contemplated, but that the disclosure will include all examples falling within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:1 . A method for optimizing a radio access network (RAN), the method comprising the steps of: a. generating, by a heuristic server unit, a heuristic recommendation for optimizing network configurations with a heuristic-based algorithm using data from the RAN as input; b. applying, by an Operations, Administration, and Maintenance (OAM) unit, the heuristic recommendation as a configuration to the RAN; c. collecting, by an AI / ML server unit, feedback data from the RAN after applying the heuristic recommendation; d. using, by the AI / ML server unit, the collected feedback data to train an AI / MLRL-based model; e. generating, by the AI / ML server unit, an RL recommendation for optimizing network configurations of the RAN with the AI / ML RL-based model using data from the RAN as input; f. comparing, by a comparison unit, the heuristic recommendation to the RL recommendation to determine whether a convergence is achieved; and g. one of:(i) switching, by a switching unit, to RL recommendations if convergence is achieved and applying future RL recommendations as a configuration to the RAN, or(ii) if convergence has not been achieved in step f, repeating steps a to f until convergence is achieved.
2. The method of claim 1 , further comprising alternatingly repeating both steps a and b, and then step e.
3. The method of claim 1 , further comprising performing both steps a and b and then step e at a ratio based on a behavioral policy determined by the heuristic-based algorithm of the heuristic server unit to make changes to the network functions of the RAN using the OAM unit and collecting feedback data from the RAN to provide as input to the heuristic server unit and the AI / ML server unit.
4. The method of claim 1 , wherein convergence is achieved when the most recent heuristic recommendations are substantially the same as previously generated heuristic recommendations.
5. The method of claim 1 , wherein convergence is achieved when at least one of the following conditions is achieved: (i) when the most recent heuristic recommendations are substantially the same as previously generated heuristic recommendations, and (ii) the AI / ML is generating recommendations that are within a predetermined percentage of actions recommended by heuristics within a predetermined time period.
6. The method of claim 1 , wherein the method is performed on a live RAN deployment with the heuristic server unit and the AI / ML server unit operating simultaneously in parallel.
7. The method of claim 1 , wherein the heuristic-based algorithm determines action states with good outcomes that can be used to generate the heuristic recommendation.
8. The method of claim 7, further comprising: training the AI / ML RL-based model using all the action states that have been determined to have good outcomes.
9. The method of claim 8, further comprising: determining action states with bad outcomes; and excluding the action states with bad outcomes from being used in live deployment, while using them as negative examples to train the AI / ML RL- based model.
10. The method of claim 1 , wherein step e further comprises performing a neighbor search constrained to a radius around action states associated with good outcomes.11 . The method of claim 1 , wherein step e further comprises transfer learning.
12. The method of claim 1 , wherein step e is performed in a lab simulation before live deployment.
13. The method of claim 1 , further comprising: performing step g until a trigger condition occurs; and repeating the method.
14. The method of claim 13, wherein the trigger condition is at least one selected from the group consisting of: KPI degradation beyond a threshold, a dataset change, an environment change, a software rollout, a new device class, a new cell site, or a channel-geometry change, and wherein the switching requires persistence of the trigger for a minimum time and reverts only after a lower threshold is satisfied.
15. A system for optimizing a radio access network, comprising: a heuristic engine configured to generate behavioral policy recommendations for radio network configuration; an AI / ML reinforcement learning engine trained using feedback data collected from an application of heuristic-based recommendations on network functions of the RAN and configured to generate behavioral policy recommendations for radio network configuration; a model inferencing engine that compares an output of the reinforcement learning model with heuristic-based recommendations during real-time operations;a switch configured to switch between heuristic-based policies from the heuristic engine and RL-based exploration from the AI / ML reinforcement learning engine; and a management system configured to apply the behavioral policy recommendations as configuration for network functions of the radio access network.
16. A method for dynamically switching between an RL-based model and a heuristicbased algorithm to optimize a radio access network (RAN), comprising the steps of: a. generating, by an AI / ML server unit, an RL recommendation for optimizing network configurations of the RAN with an AI / ML RL-based model using data from the RAN as input to the model; b. monitoring, by a comparison unit, one or more system parameters related to performance of the AI / ML RL-based model to determine whether a trigger condition has occurred; and c. one of:(i) switching, by a switching unit, to generate, by a heuristic server unit, a heuristic recommendation for optimizing network configurations with a heuristic-based algorithm, and applying by an Operations, Administration, and Maintenance (OAM) unit, the heuristic recommendation as a configuration to the RAN if the trigger is determined to have occurred, or(ii) applying, by the OAM unit, an RL recommendation as a configuration to the RAN if the trigger is determined not to have occurred.
17. The method of claim 16, wherein the trigger is at least one selected from the group consisting of: system parameters exceeding predefined thresholds indicating suboptimal performance, a change in data set, a change in environment conditions, a new software deployment that changes traffic patterns, a new device introduced in the network, an additions of a new cell site, and a change in channel geometry due to new buildings.
18. The method of claim 16, further comprising, after switching to generating the heuristic recommendation in step c, monitoring for a presence of the trigger condition, and if no longer present, switching, by the switching unit, back to generating RL recommendations by the AI / ML server unit.
19. The method of claim 16, further comprising defining, by the switching unit, an upper threshold and a lower threshold for the trigger and only switching between the RL-based model and the heuristic-based algorithm when the upper threshold or lower threshold is exceeded.
20. The method of claim 16, wherein the switching in step c occurs only if the trigger is determined to have occurred and persists for a predetermined period of time.
Citation Information
Patent Citations
Procedure for optimization of self-organizing network
US20210219384A1
Methods for artificial intelligence (AI) / machine learning (ML) model switching
WO2024211555A1