determining one or more optimized channel configurations
By optimizing the channel configuration of wireless networks using digital twins and reinforcement learning algorithms, the performance degradation problem of wireless networks in high-traffic environments was solved, network performance was improved and dynamically adapted, and the user experience was enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NOKIA NETWORKS OY
- Filing Date
- 2025-11-28
- Publication Date
- 2026-05-29
Smart Images

Figure CN122120802A_ABST
Abstract
Description
Technical Field
[0001] The example embodiments below relate to wireless communication. Background Technology
[0002] With the continued growth in demand for high-capacity and reliable wireless connectivity, ensuring optimal performance of wireless networks is to be expected. Summary of the Invention
[0003] The scope of protection sought for the various exemplary embodiments is set forth in the claims. Exemplary embodiments and features (if any) described in this specification that do not fall within the scope of the claims should be interpreted as examples helpful in understanding the various embodiments.
[0004] According to a first aspect, a method is provided, the method comprising: collecting network data from a plurality of access points within a wireless network, wherein the network data includes at least one of: performance information of the plurality of access points, spatial information indicating the location of the plurality of access points, or channel information indicating one or more channels used by the plurality of access points; selecting a digital twin method from at least two predefined digital twin methods for simulating a wireless network based on the age of the network data, wherein the at least two predefined digital twin methods include at least: a first digital twin method for simulating a wireless network using digital twins representing the plurality of access points, and a second digital twin method for simulating a wireless network using multiple digital twin instances representing the plurality of access points and one or more other access points; determining one or more optimized channel configurations for the plurality of access points based at least on the network data and the selected digital twin method; and applying one of the one or more optimized channel configurations to the plurality of access points within the wireless network.
[0005] According to the second aspect, a method of the first aspect is provided, wherein a first digital twin method is selected based on determining that the age of the network data is less than or equal to a threshold, or wherein a second digital twin method is selected based on determining that the age of the network data is greater than a threshold.
[0006] According to the third aspect, a method of the first or second aspect is provided, wherein one or more optimized channel configurations are further determined based on channel pollution information indicating: channel interference from one or more adjacent channels on one or more channels used by multiple access points, and channel occupancy, which is measured as signal power observed by at least one of the multiple access points on one or more channels.
[0007] According to the fourth aspect, a method for the third aspect is provided, which further includes: collecting channel pollution information from multiple access points; and determining, based on selecting a first digital twin method, using the channel pollution information collected from the multiple access points to determine one or more optimized channel configurations.
[0008] According to the fifth aspect, a method for the third aspect is provided, which further includes: generating channel pollution information by simulating a wireless network using multiple digital twin instances representing multiple access points and one or more other access points, based on the selection of a second digital twin method.
[0009] According to the sixth aspect, a method is provided for any of the first to fifth aspects, wherein one or more optimized channel configurations include: a plurality of optimized channel configurations determined for at least one of a plurality of access points based on the second digital twin method, the plurality of optimized channel configurations including: a channel configuration for each of the plurality of digital twin instances, wherein the method further includes: selecting a channel configuration to be applied to at least one access point from the plurality of optimized channel configurations, wherein the selection is based on the correlation between the plurality of channel configurations and one or more channels expected to be selected at at least one access point.
[0010] According to the seventh aspect, a method is provided for any of the first to sixth aspects, wherein one or more optimized channel configurations are determined by using a reinforcement learning algorithm configured to: determine a set of channel selection actions for multiple access points; apply the set of channel selection actions to a selected digital twin method; and determine one or more optimized channel configurations based on one or more rewards, wherein the one or more rewards are generated from each of the set of channel selection actions applied to the selected digital twin method, wherein the one or more rewards are at least related to performance information.
[0011] According to the eighth aspect, the method of the seventh aspect is provided, wherein the reinforcement learning algorithm is a Q-learning algorithm, which is configured to follow an epsilon-greedy policy for determining one or more optimized channel configurations.
[0012] According to the ninth aspect, a method of the seventh or eighth aspect is provided, wherein one or more rewards include: a weighted combination of at least the following reward values for each action of the set of channel selection actions: a global performance-based reward value common to multiple access points, wherein the global performance-based reward value is based on simulation results obtained from the selected digital twin method; an individual performance-based reward value for each of the multiple access points, wherein the individual performance-based reward value is based on simulation results obtained from the selected digital twin method; a channel pollution-based reward value for each of the multiple access points, wherein the channel pollution-based reward value is related to channel pollution observed by the multiple access points; and a channel reward value for incentivizing the use of one or more predefined channels.
[0013] According to the tenth aspect, the method of the ninth aspect is provided, wherein the reward value based on global performance is based on: the average network throughput value for multiple access points according to the simulation results, and the average network air time value for multiple access points according to the simulation results, and wherein the reward value based on individual performance is based on: the number of times a single access point has exceeded a predefined throughput range according to the simulation results, and the number of times a single access point has exceeded a predefined air time range according to the simulation results.
[0014] According to the eleventh aspect, a method is provided for any of the first to tenth aspects, wherein the performance information includes at least one of the following: throughput information, congestion information, or one or more received signal strength indicator values.
[0015] According to the twelfth aspect, an apparatus is provided, comprising: components for performing the methods of any of the first to eleventh aspects.
[0016] According to a thirteenth aspect, an apparatus is provided, comprising: at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform a method according to any of the first to eleventh aspects.
[0017] According to the fourteenth aspect, a computer program is provided, the computer program including instructions that, when executed by a device, cause the device to perform the method according to any of the first to eleventh aspects.
[0018] According to the fifteenth aspect, a computer-readable medium is provided, the computer-readable medium including program instructions that, when executed by a device, cause the device to perform the method according to any of the first to eleventh aspects.
[0019] According to the sixteenth aspect, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium including program instructions that, when executed by a device, cause the device to perform the method according to any of the first to eleventh aspects. Attached Figure Description
[0020] In the following description, various exemplary embodiments will be described in more detail with reference to the accompanying drawings, in which:
[0021] Figure 1 An example of the system is shown;
[0022] Figure 2 A flowchart is shown;
[0023] Figure 3 An example of a Q table is shown;
[0024] Figure 4 A flowchart is shown; and
[0025] Figure 5 An example of the device is shown. Detailed Implementation
[0026] The following embodiments are exemplary. Although the specification may refer to "a," "an," or "some" (or more) embodiments in multiple places in the text, this does not necessarily mean that every reference is for the same (or more) embodiments, or that a particular feature applies only to a single embodiment. Individual features of different embodiments may also be combined to provide other embodiments within the scope of the claims. Furthermore, the words "comprising" and "including" should be understood not to limit the described embodiments to consisting only of those features already mentioned, and such embodiments may also include features not specifically mentioned. Reference numerals in the specification and / or claims are used to illustrate embodiments with reference to the accompanying drawings, and not to limit the embodiments to these examples.
[0027] A Wireless Local Area Network (WLAN) is a type of wireless network based on the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard. WLANs use wireless communication (i.e., wireless signals) to connect devices within a limited area, such as a home, school, or office building. This allows users to maintain a network connection while moving within the coverage area. A WLAN may include one or more access points (APs) for connecting devices to the Internet.
[0028] An access point is a device that allows terminal devices (such as smartphones, tablets, or printers) to connect to a wireless network. The access point acts as a central transmitter and receiver of wireless signals, facilitating communication between the device and the internet.
[0029] WLAN dominates the supply of wireless internet access. WLAN access points are widely and densely deployed in public places, offices, and residential areas. However, given the significant increase in the amount of data traffic transmitted to and received from the internet, ensuring that deployed WLANs have sufficient capacity to meet the ever-growing data traffic demands is a challenge. For example, mobility restrictions and the rise of remote working have led to increased use and higher demand for communication networks such as indoor WLANs.
[0030] A significant proportion of broadband customer complaints relate to communication quality over WLAN. Many of these complaints, and the customer perception, can be improved through proper WLAN configuration and parameterization. Configuration and parameterization can vary over time based on the network's internal conditions and surrounding environment. For example, these changes can occur in networks already deployed by the user themselves.
[0031] Network operators can perform network optimization based on metrics that reflect the network's state at lower levels within the overall communications framework, such as data rates or network-level performance. The fundamental principle behind this optimization is improving user experience. One way to verify network optimization is by recording a decrease in customer calls to service centers. However, sometimes customer calls are not recorded, and optimization metrics can be based not only on these metrics but also on how far the user base is from falling into poor performance conditions.
[0032] Therefore, a better approach to detecting performance risks is to analyze the number of outliers in the network. Some services and applications can provide rich georeferenced data, user preferences, and terminal measurement information, generating large datasets that can be used for decision-making processes, spectrum usage improvement mechanisms, and WLAN interference mitigation (both between the same network element and between adjacent networks). Large amounts of available data cannot be easily manipulated manually. Big data and artificial intelligence (AI) mechanisms can be used to process raw data to find patterns, variables, and behaviors that are difficult to obtain using other technologies (e.g., manually). This can improve connectivity for different users, which is beneficial for the community's leisure, work, and educational activities.
[0033] The increasing complexity of WLANs, coupled with inconsistent deployments, distributed management, and network density, could negatively impact the operation of future 802.11 networks. Furthermore, improving network performance in terms of overall performance and outliers would be beneficial. Therefore, there is a need for an optimized system that can account for potential network evolution over longer periods and respond quickly. The optimized system should also be able to adjust the metrics of the specific network being optimized. However, currently, there is no optimized system available to produce such optimizations.
[0034] Some example embodiments provide a method and system for optimizing channel configuration for access points within a wireless network (e.g., a WLAN) using a digital twin based on a wireless network. A digital twin of a wireless network is a virtual model that replicates the architecture, components, and state of an actual wireless network. A digital twin can mirror the network's configuration and behavior, allowing for the analysis, optimization, and prediction of network performance in various scenarios.
[0035] It should be noted that WLAN is just one example of a wireless network, and this embodiment is not limited to WLAN. Those skilled in the art can also apply this solution to any other wireless communication network or system where control is not centralized and / or resources are shared between different networks.
[0036] A channel is a medium used for communication between an access point and connected terminal devices. It includes an identifier (i.e., a descriptor), transmission power, center frequency, and a specific frequency range within the radio spectrum used for communication between the access point and connected terminal devices. Each channel is a separate path capable of carrying data. For example, in a WLAN, channels can operate in the 2.4 GHz and 5 GHz bands.
[0037] Channel configuration refers to the specific settings and parameters assigned to a wireless communication channel used by (or multiple) access points within a wireless network. For example, channel configuration can indicate channel frequency, bandwidth, and possibly other related settings to optimize network performance and minimize interference. Proper channel configuration helps ensure efficient use of available spectrum, reduce congestion, and improve overall network reliability and speed.
[0038] For example, some example embodiments may apply machine learning algorithms (e.g., reinforcement learning algorithms) to the digital twin to optimize channel configuration. Alternatively, any other suitable technique (e.g., non-machine learning-based algorithms) may be used to optimize channel configuration.
[0039] Since the deployment of the first generation of wireless communication systems, wireless technology has continuously evolved from supporting basic coverage to meeting more advanced needs. Machine learning, a key technology for artificial intelligence, can solve complex problems without explicit programming.
[0040] However, relying on machine learning in wireless network environments without the support of digital twins presents significant challenges. Testing directly on real operational networks can lead to performance degradation, as real-time tuning and training can interfere with quality of service. Therefore, given the inherent complexity and dynamism of these infrastructures, incorporating digital twins into machine learning-based wireless network optimization systems may be beneficial.
[0041] One advantage of machine learning is its ability to learn useful information from input data, which can help improve network performance. For example, spatial and sequence features can be extracted from Received Signal Strength Indicator (RSSI).
[0042] Furthermore, machine learning-based resource, network, and mobility management algorithms are well-suited to dynamic environments. Machine learning algorithms can achieve similar performance to traditional optimization algorithms, but with significantly lower complexity, enabling rapid responses to environmental changes. Additionally, reinforcement learning can achieve fast network control based on learned policies.
[0043] Furthermore, machine learning helps achieve the goal of network self-organization. Multi-agent reinforcement learning can be used, where each node in the network can self-optimize its transmission power, channel allocation, etc.
[0044] Furthermore, by involving transfer learning, machine learning has the ability to quickly solve new problems. In wireless communication systems, there are spatiotemporal correlations, such as traffic load between adjacent areas. Therefore, it is possible to transfer knowledge acquired in one task to another related task, which can accelerate the learning process for new tasks.
[0045] Reinforcement learning may be well-suited for addressing specific challenges associated with the optimization and management of wireless networks. Wireless networks are highly dynamic and changing environments due to device mobility, interference, congestion, and other factors. Reinforcement learning is a powerful tool for adapting to these changes and continuously learning, allowing networks to automatically optimize themselves based on changing conditions. Furthermore, reinforcement learning algorithms can learn decision-making strategies based on specific metrics to maximize network performance and contribute to overall network performance improvements. Reinforcement learning algorithms are highly adaptable to different network environments and scenarios because they can continuously adjust their decisions and strategies to adapt to new network configurations or changes in traffic load or Quality of Service (QoS) requirements.
[0046] A key element of some reinforcement learning systems is the environment model. This is something that mimics the behavior of the environment, or more generally, allows for inferences about how the environment will behave. For example, given a state and an action, the model can predict the next outcome state and the next reward. Models can be used for planning, that is, deciding on course of action by considering possible future scenarios before they occur. Techniques for solving reinforcement learning problems using models and planning are called model-based techniques, as opposed to the simpler model-free techniques of explicit trial-and-error learning. In other words, the scope of reinforcement learning can range from low-level trial-and-error learning to high-level deliberate planning.
[0047] The continuous dynamics of networks may necessitate real-time adaptability. Digital twins, by virtually replicating the network environment, provide an accurate and up-to-date representation for immediate adjustments. Furthermore, the introduction of digital twins enables optimization with less intrusion into the real network, as testing and adjustments can be performed virtually before deployment to the physical environment, minimizing potential adverse effects. Similarly, to support reinforcement learning, digital twins provide a safe and controlled environment for comprehensive testing, accelerating the learning process and contributing to greater convergence of the optimized system. In conclusion, the combination of machine learning and digital twins not only improves operational efficiency but also allows for addressing the dynamic challenges of wireless networks with greater agility and accuracy.
[0048] Figure 1 An example embodiment of a system for optimizing a wireless network by utilizing a digital twin approach is shown. For example, the optimization system 120 may reside in a cloud computing platform, or in... Figure 5 In the device 500 shown.
[0049] refer to Figure 1 The optimization system 120 collects network data from multiple access points 111, 112, 113 (or a cluster of access points) within a wireless network (NW) 110 (such as a wireless local area network). For example, the network data may include performance information (or performance metrics) of the multiple access points 111, 112, 113. The multiple access points 111, 112, 113 may periodically report performance information to the optimization system 120.
[0050] Performance information refers to data indicating how well a given access point 111, 112, or 113 is performing. For example, performance information may include at least one of the following: throughput information, congestion information, or one or more Received Signal Strength Indicator (RSSI) values (e.g., measured by multiple access points 111, 112, or 113). Throughput refers to the amount of data successfully transmitted over the network within a given time period. Congestion refers to an overload condition where the demand for data transmission for traffic exceeds the channel capacity, thus slowing down or even stopping actual data transmission. The RSSI value indicates the power level of the received radio signal.
[0051] Network data may also include network details reported by multiple access points 111, 112, 113, such as spatial information indicating the location of multiple access points 111, 112, 113, and / or channel information indicating one or more channels used by multiple access points 111, 112, 113.
[0052] Location information refers to data (e.g., geographic or spatial coordinates) indicating the physical location of access points 111, 112, and 113 within wireless network 110. Location information can be beneficial for optimizing network performance because it helps in understanding network layout, planning coverage areas, and managing interference and signal strength.
[0053] Channel information refers to data related to a specific frequency channel used by access points 111, 112, and 113 within wireless network 110. Channel information can be beneficial for optimizing wireless network 110 because it helps in selecting the optimal channel to minimize interference and maximize performance. For example, channel information may include at least one of the following details: channel number, frequency, bandwidth, or channel utilization. Channel number refers to the specific channel being used (e.g., channel 1, channel 6, or channel 11 in the 2.4 GHz band). Frequency refers to the frequency range of the channel (e.g., channel 1 is 2.412 GHz). Bandwidth refers to the width of the channel. Channel utilization refers to the degree of channel usage (e.g., metrics including occupancy and interference levels).
[0054] Multiple access points 111, 112, and 113 can refer to access points controlled by the optimization system 120 and belonging to a communication service provider (CSP). Multiple access points 111, 112, and 113 can also be referred to as controlled access points. Although Figure 1 The image shows three access points: 111, 112, and 113. However, it should be noted that the number of access points can be more than three. In other words, multiple access points can include two or more access points.
[0055] Multiple access points 111, 112, and 113 may experience interference from one or more other access points 114 and 115. One or more other access points 114 and 115 can refer to access points within wireless network 110 that are not controlled by optimization system 120 and from which network data is not collected. Alternatively, one or more other access points 114 and 115 may be part of a separate (adjacent) wireless network. One or more other access points 114 and 115 may belong to one or more different CSPs than the CSP controlling multiple access points 111, 112, and 113. Alternatively, one or more other access points 114 and 115 may belong to the same CSP controlling multiple access points 111, 112, and 113. One or more other access points 114 and 115 can be identified by at least one (controlled) access point 111, 112, and 113 reporting their presence. One or more other access points 114 and 115 may also be referred to as invasive access points or uncontrolled access points. Although Figure 1 The image shows two uncontrolled access points (external access points) 114 and 115, but it should be noted that the number of uncontrolled access points may also be different than two (i.e., one or more).
[0056] The optimization system 120 can also collect channel pollution information (or statistics) from multiple (controlled) access points 111, 112, 113, wherein the channel pollution information indicates: 1) channel interference from one or more adjacent (nearby) channels on one or more channels used by the multiple (controlled) access points 111, 112, 113; and 2) channel occupancy observed by at least one (controlled) access point among the multiple (controlled) access points 111, 112, 113 on one or more channels. The channel pollution information may also be referred to as an external channel pollution mask.
[0057] Channel occupancy is measured as the signal power (or received signal strength) caused by any other access point(s) (e.g., other controlled access points 111, 112, 113 and / or uncontrolled access points 114, 115), as observed by at least one of the (controlled) access points 111, 112, 113 on one or more channels. The effect of channel occupancy is primarily contention, but the channel can still be used by sharing airtime among different access points. Airtime refers to the amount of time a given access point occupies a communication channel to send or receive data.
[0058] Channel interference refers to the interruption or degradation of wireless communication signals caused by overlapping or competing signals on adjacent (nearby) channels (e.g., signal leakage from adjacent channels). In wireless networks, multiple access points may operate on the same frequency band, leading to interference. This interference can cause network performance degradation, reduced data throughput, increased latency, and / or connection instability. Typically, channel interference may be smaller in magnitude than channel occupancy.
[0059] The optimization system 120 processes the collected network data to determine the age of the network data. The age of the network data refers to the time elapsed since the network data was reported from multiple access points 111, 112, and 113. Each access point 111, 112, and 113 can generate a timestamp, which is inserted into the reports sent by that access point, indicating the date (e.g., year, month, day) and / or time (e.g., hour, minute, second, millisecond) when the timestamp was generated. Different access points 111, 112, and 113 can send reports at different timestamps, but at regular time intervals (e.g., every 15 minutes or every 24 hours). For example, if reports are generated once a day during peak hours, the timestamp may only indicate the date, not the time (in which case, timestamps generated by different access points on the same day would be the same). The optimization system 120 can compare the timestamps with the current date and / or time (e.g., indicated by the internal clock of the optimization system 120 or the cloud computing platform) to determine the age of the network data. For example, the age of network data can be measured as the number of days since the report was generated relative to the current date (a small portion of a day can also be considered a day). If the network data includes multiple different timestamps (e.g., from different dates), the age of the network data can be determined, for example, by comparing the latest timestamp (from the last received report) with the current date and / or time, or by comparing the oldest timestamp (from the first received report) with the current date and / or time, or by comparing the average or median of the timestamps with the current date and / or time.
[0060] Based on the age of the collected network data, the optimization system 120 selects one of at least two predefined digital twin methods 131 and 132 to simulate the wireless network 110 in the network simulator 130. The network simulator 130 can also run on a cloud computing platform or device 500.
[0061] At least two predefined digital twin methods include at least: a first digital twin method for simulating a wireless network using digital twins representing multiple (controlled) access points 111, 112, 113; and a second digital twin method for simulating a wireless network 110 using multiple digital twin instances representing multiple (controlled) access points 111, 112, 113, and one or more other (uncontrolled) access points 114, 115.
[0062] A digital twin instance refers to a specific version or copy of a digital twin. In the context of a wireless network, multiple digital twin instances can be created to represent different scenarios, configurations, or conditions of the wireless network 110. In other words, a digital twin refers to an overall virtual model, and digital twin instances are individual copies used to explore and test different scenarios within that model.
[0063] The terms “first digital twin approach” and “second digital twin approach” used in this article are used to distinguish digital twin approaches, and they do not necessarily refer to any particular order of digital twin approaches.
[0064] The first digital twin method 131 can also be referred to as a digital twin without external access points because it is a copy of the wireless network 110 that does not consider one or more other (uncontrolled) access points 114, 115. The first digital twin method can be used with up-to-date hotspot traffic. The first digital twin method 131 can be selected if the network data is recent (i.e., not too old). For example, the first digital twin method 131 can be selected based on determining that the age of the network data is less than or equal to a threshold (e.g., up to five days). It should be noted that five days is just one example of a threshold, and any other desired time period can alternatively be used as the threshold.
[0065] On the other hand, the second digital twin method 132 also considers one or more other (uncontrolled) access points 114, 115 (e.g., with hourly traffic profiles). The second digital twin method 132 can be selected if current network data is unavailable. For example, the second digital twin method 132 can be selected based on determining that the age of the network data is above a threshold (e.g., more than five days). Compared to the first digital twin method, the second digital twin method may require more computation time because multiple digital twin instances are simulated. When the latest network data is unavailable, multiple digital twin instances may be needed to obtain a more accurate approximation of the real network 110 (i.e., in this case, a single digital twin may not be sufficient).
[0066] exist Figure 1In this context, the digital twin details associated with the first digital twin method 131 and the digital twin instance details associated with the second digital twin method 132 refer to the attributes of different elements in the wireless network 110, such as the location (e.g., X, Y, Z coordinates) of controlled access points 111, 112, 113 and / or uncontrolled access points 114, 115, channel settings, radio attenuation between different devices, traffic models used, etc. Digital twin metrics and digital twin instance metrics refer to the output values from the digital twin used for optimization, such as the average throughput of different network elements (e.g., access points 111, 112, 113) and / or the maximum capacity of any network element, the minimum capacity of any network element, congestion, available airtime, used airtime, noise, and interference, etc.
[0067] The optimization system 120 determines one or more optimized channel configurations for a plurality of (controlled) access points 111, 112, 113, based at least on network data and a selected digital twin method. In other words, one or more optimized channel configurations may be determined for each access point, or for at least one of the plurality of access points 111, 112, 113.
[0068] There may be a feedback loop between the optimization system 120 and the network simulator 130, which allows the optimization system 120 to iteratively update and refine the channel configuration based on the simulation results until one or more optimized channel configurations are obtained.
[0069] Channel pollution information can also be used to determine one or more optimized channel configurations. For example, based on selecting a first digital twin method (i.e., if the network data is up-to-date), the optimization system 120 can determine one or more optimized channel configurations using channel pollution information collected from multiple access points 111, 112, 113.
[0070] Alternatively, based on the selection of a second digital twin method (i.e., if the network data is outdated), channel pollution information can be generated by simulating wireless network 110 using multiple digital twin instances representing multiple access points 111, 112, 113, and one or more other access points 114, 115. In other words, in this case, channel pollution information (or statistics) can be collected from the digital twin instances (rather than from the real access points). The channel pollution information collected from the digital twin instances can also be referred to as a digital twin (DT) channel foreign pollution mask.
[0071] For example, one or more optimized channel configurations can be determined using a reinforcement learning algorithm configured to: determine a set of channel selection actions for multiple access points; apply the set of channel selection actions to a selected digital twin method; and determine one or more optimized channel configurations based on one or more rewards, wherein the one or more rewards are generated from each of the set of channel selection actions applied to the selected digital twin method.
[0072] Channel selection action means that the reinforcement learning algorithm selects a channel from the set of candidate channels in each of the multiple access points 111, 112, 113 within the network simulator 130 (e.g., following an epsilon-greedy policy). In other words, the selected action is the action with respect to each (virtual) AP, where it follows an action policy to change the channel within the frequency band.
[0073] The optimization system 120 can use the collected network data to define one or more rewards and different weights for the reinforcement learning algorithm, depending on the desired type of optimization. Alternatively, at least one subset of rewards and / or weights can be predefined or user-defined (e.g., via a graphical user interface). For example, one or more rewards can be at least related to performance information, which can be included in the network data. One or more rewards can also be related to channel pollution information (i.e., channel pollution information can be used to derive some rewards(s)). One or more rewards can be defined to improve performance and reduce the number of outliers, such that the action selected by a given (virtual) access point considers not only improving its performance but also reducing the number of outliers in the wireless network 110. Weights can be used to add more weight or relevance to some rewards(s) compared to others (e.g., a higher weight can be added to throughput than to congestion).
[0074] Actions selected by reinforcement learning algorithms are applied to virtual representations of access points 111, 112, and 113 in network simulator 130 to test various channel configurations using digital twins. The digital twins also generate their own performance metrics (i.e., simulation results) and / or network details via network simulator 130.
[0075] One or more rewards are applied to these actions using simulation results reported by the selected digital twin method (and possibly by channel external contamination masking), thereby generating a feedback loop between the optimized system 120 and the digital twin, which produces one or more optimized channel configurations. Upon algorithm convergence (i.e., to the real network), the channel configuration that provides the highest reward (multiple) can be applied to multiple access points 111, 112, 113 (each of the multiple access points 111, 112, 113). In other words, individual channel configurations can be determined and applied individually for each of the multiple access points 111, 112, 113.
[0076] When a second digital twin method with multiple digital twin instances is selected, optimized channel configurations can be received for each digital twin instance, resulting in multiple channel configurations for each access point 111, 112, 113. In this case, a configuration selector can be used to select a single channel configuration to be applied to a given access point.
[0077] In other words, one or more optimized channel configurations may include: multiple optimized channel configurations determined based on the second digital twin method for at least one of the multiple access points 111, 112, 113, wherein the multiple optimized channel configurations include: one channel configuration for each of the multiple digital twin instances. A channel configuration to be applied to at least one access point can be selected from the multiple optimized channel configurations, wherein this selection can be based on the correlation (or comparison) between the multiple channel configurations and one or more channels expected to be selected at at least one access point.
[0078] "One or more channels expected to be selected at at least one access point" refers to the competitive local channel configuration at at least one access point. Typically, access points operate autonomously and select the locally optimal channel configuration. However, the optimization system 120 can impose different channel configurations that are optimal for a larger area than a single access point. Correlation ensures that the deviation between the autonomous local configuration and the centrally imposed configuration is not too large. For example, the configuration selector can select the channel configuration that most closely matches the most likely local channel configuration (i.e., the channel most likely to be selected at each AP based on channel information) from multiple optimized channel configurations. In other words, the optimization system 120 can be configured to preferentially select a channel configuration that minimizes the number of channel variations at the AP relative to the one or more channels originally used at the AP.
[0079] The optimization system 120 applies the selected channel configuration to multiple (controlled) access points 111, 112, 113 (each of the multiple (controlled) access points 111, 112, 113) within the wireless network 110 to enhance network performance. In other words, individual channel configurations can be selected and applied to each actual access point 111, 112, 113.
[0080] Figure 2 A flowchart is shown according to an example embodiment of a method for optimizing a wireless network by utilizing a digital twin approach. Figure 2 The method can be implemented in a cloud computing platform or by... Figure 5 The illustrated device 500 performs this function. For example, device 500 may be, include, or be incorporated into a cloud server or any other computing device.
[0081] refer to Figure 2 In block 201, the optimization system 120 periodically collects reports, metrics, and network details data (e.g., the network data mentioned above) generated by multiple access points 111, 112, and 113 within the wireless network 110.
[0082] In block 202, network data is processed and the age of the network data is determined. Based on the age of the network data, a digital twin method is selected from at least two predefined digital twin methods 131 and 132 to simulate the wireless network 110 in the network simulator 130.
[0083] At least two predefined digital twin methods include at least: a first digital twin method 131 for simulating a wireless network 110 using digital twins representing multiple access points 111, 112, 113; and a second digital twin method 132 for simulating a wireless network 110 using multiple digital twin instances representing multiple access points 111, 112, 113, and one or more other access points 114, 115 from which network data is not collected.
[0084] In block 203, based on network data, one or more rewards used in the optimization are defined and integrated, and specific optimization weights are selected.
[0085] In block 204, one or more rewards are applied to the action using simulation results reported by the selected digital twin method. As an example, the following three types of simulator data can be used to apply one or more rewards: channel contamination information (channel external contamination mask), performance information, and network details (e.g., location information and channel information).
[0086] If the age of the network data collected from the real network 110 is up-to-date (e.g., below a threshold), then channel pollution information (channel foreign pollution mask) of the real wireless network 110 is provided, and updated indicators and reports of the foreign-free digital twin (i.e., the first digital twin method 131) are generated by the network simulator 130.
[0087] Alternatively, if the network data collected from the real network 110 is outdated (e.g., above a threshold), a second digital twin method 132 is followed to generate multiple digital twin instances using one or more other (uncontrolled) access points 114, 115 (i.e., outsiders) to obtain a more accurate approximation of the optimal solution. In this case, the network simulator 130 generates channel pollution information (channel outside pollution mask), as well as reports and indicators, for each digital twin instance.
[0088] In block 205, after the application of one or more rewards, an action policy is followed for each AP 111, 112, 113 to explore or develop channels in the frequency band by selecting specific channels and applying them to the corresponding digital twins. In this way, a loop is generated until the algorithm converges to obtain one or more optimized channel configurations, and then one of the one or more optimized channel configurations is applied to each of the multiple access points 111, 112, 113 within the real wireless network 110. For example, the one or more optimized channel configurations can be determined using a reinforcement learning algorithm as described above.
[0089] In one embodiment, the reinforcement learning algorithm may be a Q-learning algorithm configured to follow an epsilon-greedy policy for determining one or more optimized channel configurations.
[0090] Epsilon-greedy strategies are techniques used in reinforcement learning to balance exploration and exploitation when making decisions. Exploration means trying new actions to discover their effects and potentially find better options. Exploration means choosing the most known action based on past experience to maximize reward. In an epsilon-greedy strategy, a parameter called epsilon(ϵ) determines the balance between exploration and exploitation. With probability ϵ, the algorithm chooses a random action (exploration), and with probability 1-ϵ, it chooses the most known action (exploration). ϵ can start with a high value to encourage exploration and gradually decrease over time to support exploitation. This strategy helps the algorithm learn optimal actions while still considering new possibilities.
[0091] Q-learning methods in reinforcement learning offer good scalability and computational efficiency. For example, it may be suitable for applications that require real-time changes. Q-learning algorithms can handle stochastic environments while maintaining up-to-date estimates. Within Q-learning, appropriate policies can be established to strike an optimal balance between exploration and exploitation of actions in a given environment. For this reason, epsilon-greedy policies can be beneficial, as they guarantee continuous exploration when new actions are proposed. Furthermore, epsilon-greedy policies do not require complex probability distribution calculations or uncertainty-based computations and allow for adjusting the epsilon value to control the balance between exploration and exploitation. This provides flexibility to adapt to different scenarios and system preferences.
[0092] The Q-learning algorithm operates using multiple parameters. On one hand, there exists an agent performing actions within an execution environment (e.g., wireless network 110). Then, the environment returns a series of rewards, which are stored in a table called the Q-table, along with states that form a new layout for the environment.
[0093] Figure 3 An example of Q-table 300 is shown. Q-table 300 is formed by the state 301 of the environment and the possible actions 302 to be performed. In this case, the agent is a (virtual) AP performing the channel selection action, and the state is a different (virtual) AP of the selected cluster to be optimized. Therefore, the state is: And the action is:
[0094] To obtain the values of the Q table, one or more custom rewards supported by the digital twin can be defined. For example, one or more rewards may include: a weighted combination of at least the following reward values for each action in the set of channel selection actions: a global performance-based reward value shared by multiple access points 111, 112, and 113 (denoted as global). reward ); the individual performance-based reward value (represented as indiv) for each of the multiple access points 111, 112, and 113. reward ); the channel pollution-based reward value (denoted as aliens) for each of the multiple access points 111, 112, and 113. reward ); and the channel reward value (denoted as key) used to incentivize the use of one or more predefined channels. channel ).
[0095] The reward value based on global performance and the reward value based on individual performance can be based on simulation results obtained from the selected digital twin method. For example, the reward value based on global performance can be based on: the average network throughput value for multiple access points 111, 112, and 113 according to the simulation results, and the average network air time value for multiple access points 111, 112, and 113 according to the simulation results.
[0096] For example, reward values based on individual performance could be based on: the number of times a single access point has exceeded a predefined throughput range according to simulation results, and the number of times a single access point has exceeded a predefined airtime range according to simulation results.
[0097] The reward value based on channel pollution is related to the channel pollution observed through multiple access points 111, 112, and 113 indicated by channel pollution information.
[0098] These reward values will be described in more detail below.
[0099] The actions performed by each (virtual) AP can follow an epsilon-greedy policy, where they are based on The value is used to select a random action or the best possible action: With probability of In the case of With probability of In the case of in:
[0100] When a (virtual) AP selects a given channel, the wireless network 110 with that new state is simulated in a digital twin. Network simulator 130 returns the optimal values for throughput and airtime for each AP. Using these normalized throughput and airtime values, global... reward It can be obtained, for example, as follows: in: Average network throughput normalized per AP ≡Average network time of implementation per AP (normalized) ≡Based on average network throughput per AP ≡Based on minimum network throughput per AP ≡Based on maximum network throughput per AP ≡Average network time per AP ≡Airtime per AP minimum network implementation ≡Airtime per AP maximum network implementation ≡ Weights assigned to average network throughput ≡Weights assigned to average network time
[0101] In addition, individual rewards per AP can also be considered, for example, by obtaining reference values for the airtime and throughput of wireless network 110 with its initial configuration (before optimization). APs with throughput below the reference value and APs with airtime above the reference value can be penalized to obtain an individual reward. reward : in: Number of APs outside the normalized throughput range ≡Number of APs outside the normalized spacetime range ≡Number of APs outside the throughput range Minimum number of APs outside the throughput range ≡Maximum number of APs outside the throughput range ≡Number of APs outside the airtime range Minimum number of APs outside the airtime range ≡Maximum number of APs outside the airtime range ≡ Weights assigned to individual network throughput ≡Weights assigned to individual network time in the air
[0102] In addition to these merit values, channel pollution (i.e., channel interference and channel occupancy) introduced by one or more other (uncontrolled) access points 114, 115 by each controlled AP 111, 112, 113 can also be considered. Channel pollution can be indicated by channel pollution information (channel external pollution mask) of wireless network 110. For example, if the difference between the RSSI of an AP from its site (STA) and the channel pollution observed or received by the AP from one or more other (uncontrolled) APs 114, 115 in a given channel exceeds 10 dB, then a reward value (aliens) based on channel pollution is awarded. reward This can be set to a maximum value (e.g., a value of 1). A station (STA) refers to any device (e.g., a laptop, smartphone, tablet, printer) that has the ability to connect to a wireless network.
[0103] Therefore, for example, aliens reward It can be defined as: in: ≡Based on average external RSSI per channel ≡Average RSSI of sites in AP ≡Minimum RSSU difference between AP RSSI and external RSSI ≡ Maximum RSSU difference between AP RSSI and external RSSI
[0104] In addition, channel rewards can be added to one or more predefined channels (e.g., channels 1, 6, and 11, as these are the only channels in the 2.4 GHz spectrum without frequency overlap). These channel rewards can incentivize APs to assign specific priorities to one or more predefined channels.
[0105] The rewards described above can be weighted and combined to form a total reward, as shown below: in:
[0106] Different optimization methods can be pursued by defining rewards in the optimization algorithm. For example, these methods can include at least three different options: a first method where only channels 1, 6, and 11 are used during optimization; a second method where channels 1, 6, and 11 have some priority, but other channels are also allowed to be used; and a third method where any channel (e.g., channels 1 through 11) can be used without any restrictions. The channel range optimization method for these target APs can be defined by changing the reward weights, for example, as follows:
[0107] This total reward, together with the learning rate (e.g., α=0.1) and the discount factor (e.g., γ=0.9) (whose value can be obtained experimentally), forms the reward function:
[0108] make In the state The actions taken in, and In the state The Q-learning algorithm converges if the following conditions are met regarding the rewards received:
[0109] Figure 4 A flowchart is shown according to an example embodiment of a method for optimizing a wireless network by utilizing a digital twin approach. Figure 4 The method can be implemented in a cloud computing platform (e.g., by optimization system 120), or by... Figure 5 The illustrated device 500 performs this function. For example, device 500 may be, include, or be incorporated into a cloud server or any other computing device.
[0110] refer to Figure 4 In block 401, network data is collected from multiple access points 111, 112, and 113 within wireless network 110, wherein the network data includes at least one of the following: performance information of the multiple access points 111, 112, and 113; spatial information indicating the location of the multiple access points 111, 112, and 113; or channel information indicating one or more channels used by the multiple access points 111, 112, and 113. Wireless network 110 may be a wireless local area network (WLAN) or any other type of wireless network.
[0111] In block 402, based on the age of the network data, a digital twin method is selected from at least two predefined digital twin methods to simulate wireless network 110.
[0112] At least two predefined digital twin methods include at least: a first digital twin method 131 for simulating a wireless network 110 using digital twins representing multiple access points 111, 112, 113; and a second digital twin method 132 for simulating a wireless network 110 using multiple digital twin instances representing multiple access points 111, 112, 113, and one or more other access points 114, 115 from which network data is not collected.
[0113] For example, the first digital twin method can be selected based on determining that the age of the network data is below or equal to a threshold.
[0114] As another example, a second digital twin approach could be selected based on determining that the age of the network data is above a threshold.
[0115] In block 403, based at least on network data and the selected digital twin method, one or more optimized channel configurations are determined for multiple access points 111, 112, 113 (each of the multiple access points 111, 112, 113).
[0116] One or more optimized channel configurations can also be determined based on channel pollution information, which indicates: channel interference from one or more adjacent channels on one or more channels used by multiple access points 111, 112, 113, and channel occupancy as measured as signal power from one or more other access points 114, 115 on one or more channels used by multiple access points 111, 112, 113.
[0117] Channel pollution information can be collected from multiple access points 111, 112, and 113 of a real wireless network 110. Based on the selection of a first digital twin method, the channel pollution information collected from the multiple access points can be used to determine one or more optimized channel configurations.
[0118] Alternatively, based on the selection of a second digital twin method, channel pollution information can be generated by simulating wireless network 110 using multiple digital twin instances representing multiple access points 111, 112, 113, and one or more other access points 114, 115.
[0119] For example, one or more optimized channel configurations can be determined using the reinforcement learning algorithm or Q-learning algorithm described above. Alternatively, one or more optimized channel configurations can be determined using another type of machine learning algorithm (e.g., supervised learning algorithm) or by using a non-machine learning-based algorithm.
[0120] In block 404, one of the optimized channel configurations is applied to multiple access points 111, 112, 113 (each of the multiple access points 111, 112, 113) within the wireless network 110.
[0121] pass Figure 2 and Figure 4 The described blocks, related functions, and information exchanges do not have an absolute temporal order, and some of them may be executed concurrently or in a different order than described. Other functions may also be executed between or within them, and may send additional information and / or apply additional rules. Some of a block or part of a block, or one or more messages, may also be omitted or replaced by the corresponding block or part of a block, or one or more messages.
[0122] As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where a list of two or more elements is connected by “and” or “or”, means at least any one of the elements, or at least any two or more of the elements, or at least all of the elements.
[0123] Figure 5 An example of apparatus 500 is shown, which includes one or more example embodiments for performing the example embodiments described above (e.g., Figure 2 or Figure 4 The device 500 is a component of the method. For example, the device 500 may be, include, or be included in a cloud server or any other computing device.
[0124] For example, device 500 may include a circuit system or chipset suitable for implementing one or more of the example embodiments described above. Device 500 may be an electronic device or computing system including one or more electronic circuit systems. Device 500 may include an optimized circuit system 510, such as at least one processor, and at least one memory 520 storing instructions 522 that, when executed by the at least one processor, cause device 500 to perform one or more of the example embodiments described above. Such instructions 522 may, for example, include computer program code (software). The at least one processor and the at least one memory storing the instructions may provide components for providing or causing execution of any of the methods and / or blocks described above.
[0125] A processor is coupled to memory 520. The processor is configured to read data from memory 520 and write data to memory 520. Memory 520 may include one or more memory cells. Memory cells may be volatile or non-volatile. It should be noted that one or more non-volatile memory cells and one or more volatile memory cells may be present, or alternatively, one or more non-volatile memory cells, or alternatively, one or more volatile memory cells. Volatile memory may be, for example, random access memory (RAM), dynamic random access memory (DRAM), or synchronous dynamic random access memory (SDRAM). Non-volatile memory may be, for example, read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), flash memory, optical storage, or magnetic storage. In general, memory may be referred to as a non-transitory computer-readable medium. As used herein, the term "non-transitory" is a limitation on the medium itself (i.e., tangible, not tactile), rather than a limitation on the persistence of data storage (e.g., RAM and ROM). Memory 520 stores computer-readable instructions that are executed by the processor. For example, non-volatile memory stores computer-readable instructions, and the processor uses volatile memory to execute instructions for temporary storage of data and / or instructions.
[0126] The computer-readable instructions may have been pre-stored in memory 520, or alternatively, they may be received by the device via an electromagnetic carrier signal, and / or copied from a physical entity (such as a computer program product). Execution of the computer-readable instructions causes the device 500 to perform one or more of the functions described above.
[0127] The memory 520 can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and / or removable memory.
[0128] The device 500 may also include or be connected to a communication interface 530, which includes hardware and / or software for establishing a communication connection according to one or more communication protocols. The communication interface 530 may include at least one transmitter (Tx) and at least one receiver (Rx) that can be integrated into or connected to the device 500. The communication interface 530 may include one or more components, such as a power amplifier, digital front-end (DFE), analog-to-digital converter (ADC), digital-to-analog converter (DAC), frequency converter, (de)modulator, and / or encoder / decoder circuitry controlled by a corresponding control unit.
[0129] Communication interface 530 provides the device with the ability to communicate with wireless network 110. For example, communication interface 530 can provide radio, cable, or fiber optic interfaces for one or more access points 111, 112, 113 of wireless network 110.
[0130] It should be noted that device 500 may also include Figure 5 Various components are not shown. These components can be hardware components and / or software components.
[0131] As used in this application, the term "circuit system" may refer to one or more or all of the following: (a) a hardware circuit implementation only (such as an implementation only in analog and / or digital circuit systems); and (b) a combination of hardware circuits and software, such as (if applicable): (i) a combination of (multiple) analog and / or digital hardware circuits having software / firmware, and (ii) any part of (multiple) hardware processors having software (including (multiple) digital signal processors, software, and (multiple) memories, which work together to enable a device (such as a mobile phone or a server) to perform various functions); and (c) (multiple) hardware circuits and / or (multiple) processors, such as (multiple) microprocessors or a portion thereof, which require software (e.g., firmware) to operate, but may be absent when operation is not required.
[0132] This definition of circuit system applies to all uses of the term in this application (including in any claim). As another example, as used in this application, the term circuit system also covers only hardware circuitry or a processor (or multiple processors) or portions of hardware circuitry or a processor and its accompanying software and / or firmware implementation. For example, and if applicable to a particular claim element, the term circuit system also covers baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices or other computing or network devices.
[0133] The techniques and methods described herein can be implemented by various means. For example, these techniques can be implemented in hardware (one or more devices), firmware (one or more devices), software (one or more modules), or combinations thereof. For hardware implementation, the apparatus(s) of the example embodiments can be implemented within one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to perform the functions described herein, or combinations thereof. For firmware or software, the implementation can be executed by modules (e.g., programs, functions, etc.) of at least one chipset that perform the functions described herein. Software code can be stored in memory units and executed by a processor. The memory units can be implemented within or outside the processor. In the latter case, it can be communicatively coupled to the processor via various means as known in the art. Additionally, the components of the systems described herein can be rearranged and / or supplemented by additional components to facilitate the implementation of the relevant aspects, etc., and they are not limited to the precise configurations illustrated in the given figures, as will be understood by those skilled in the art.
[0134] Those skilled in the art will understand that, with advancements in technology, the proposed concepts can be implemented in various ways within the scope of the claims. Embodiments are not limited to the exemplary embodiments described above, but can vary within the scope of the claims. Therefore, all words and expressions should be interpreted broadly, and they are intended to be illustrative rather than limiting.
Claims
1. A method for communication, comprising: (401) Network data is collected from multiple access points (111, 112, 113) within a wireless network (110), wherein the network data includes at least one of the following: performance information of the multiple access points (111, 112, 113), spatial information indicating the location of the multiple access points (111, 112, 113), or channel information indicating one or more channels used by the multiple access points (111, 112, 113); Based on the age of the network data, one digital twin method is selected (402) from at least two predefined digital twin methods (131, 132) used to simulate the wireless network (110). The at least two predefined digital twin methods (131, 132) mentioned above include at least: A first digital twin method (131) is used to simulate the wireless network (110) using digital twins representing the plurality of access points (111, 112, 113), and A second digital twin method (132) is used to simulate the wireless network (110) using multiple digital twin instances representing the plurality of access points (111, 112, 113) and one or more other access points (114, 115). Based at least on the network data and the selected digital twin method, determine (403) one or more optimized channel configurations for the plurality of access points (111, 112, 113); and Apply (404) one of the one or more optimized channel configurations to the plurality of access points (111, 112, 113) within the wireless network (110).
2. The method according to claim 1, wherein the first digital twin method (131) is selected based on determining that the age of the network data is less than or equal to a threshold, or The second digital twin method (132) is selected based on determining that the age of the network data is higher than the threshold.
3. The method of claim 1, wherein the one or more optimized channel configurations are further determined based on channel pollution information. The channel pollution information indicates: Channel interference from one or more adjacent channels on one or more channels used by the plurality of access points (111, 112, 113), and Channel occupancy, the credit occupancy being measured as signal power observed by at least one of the plurality of access points (111, 112, 113) on the one or more channels.
4. The method according to claim 3, further comprising: The channel pollution information is collected from the multiple access points (111, 112, 113); as well as Based on the selection of the first digital twin method, the channel pollution information collected from the plurality of access points (111, 112, 113) is used to determine the one or more optimized channel configurations.
5. The method according to claim 3, further comprising: Based on the selection of the second digital twin method, the channel pollution information is generated by simulating the wireless network (110) using the plurality of digital twin instances representing the plurality of access points (111, 112, 113) and the one or more other access points (114, 115).
6. The method according to any one of claims 1 to 5, wherein the one or more optimized channel configurations comprise: Based on the second digital twin method (132), multiple optimized channel configurations are determined for at least one of the plurality of access points (111, 112, 113), the multiple optimized channel configurations including: one channel configuration for each of the plurality of digital twin instances. The method further includes: A channel configuration to be applied to the at least one access point is selected from the plurality of optimized channel configurations, wherein the selection is based on the correlation between the plurality of channel configurations and one or more channels that are expected to be selected at the at least one access point.
7. The method according to any one of claims 1 to 5, wherein the one or more optimized channel configurations are determined by using a reinforcement learning algorithm, the reinforcement learning algorithm being configured to: Determine the set of channel selection actions for the plurality of access points (111, 112, 113); Apply the set of channel selection actions to the selected digital twin method; and The one or more optimized channel configurations are determined based on one or more rewards, wherein the one or more rewards are generated from each of the set of channel selection actions applied to the selected digital twin method. The one or more rewards mentioned therein are at least related to the performance information.
8. The method of claim 7, wherein the reinforcement learning algorithm is a Q-learning algorithm, the Q-learning algorithm being configured to follow an epsilon-greedy policy for determining the one or more optimized channel configurations.
9. The method of claim 7, wherein the one or more rewards comprise: For each of the set of channel selection actions, a weighted combination of at least the following reward values: The shared global performance-based reward value for the multiple access points (111, 112, 113), wherein the global performance-based reward value is based on simulation results obtained from the selected digital twin method. The reward value based on individual performance for each of the plurality of access points (111, 112, 113), wherein the reward value based on individual performance is based on simulation results obtained from the selected digital twin method. The reward value based on channel pollution for each of the plurality of access points (111, 112, 113), wherein the reward value based on channel pollution is related to the channel pollution observed by the plurality of access points (111, 112, 113), and Channel reward values used to incentivize the use of one or more predefined channels.
10. The method of claim 9, wherein the reward value based on global performance is based on: the average network throughput value for the plurality of access points (111, 112, 113) according to the simulation results, and the average network airtime value for the plurality of access points (111, 112, 113) according to the simulation results, and The individual performance-based reward value is based on: the number of times a single access point has exceeded a predefined throughput range according to the simulation results, and the number of times the single access point has exceeded a predefined airtime range according to the simulation results.
11. The method according to any one of claims 1 to 5, wherein the performance information includes at least one of the following: throughput information, congestion information, or one or more received signal strength indicator values.
12. A communication apparatus (500), comprising: Components for collecting network data from multiple access points (111, 112, 113) within a wireless network (110), wherein the network data includes at least one of the following: performance information of the multiple access points (111, 112, 113), spatial information indicating the location of the multiple access points (111, 112, 113), or channel information indicating one or more channels used by the multiple access points (111, 112, 113); A component for selecting a digital twin method from at least two predefined digital twin methods (131, 132) for simulating the wireless network (110) based on the age of the network data. The at least two predefined digital twin methods (131, 132) mentioned above include at least: A first digital twin method (131) is used to simulate the wireless network (110) using digital twins representing the plurality of access points (111, 112, 113), and A second digital twin method (132) is used to simulate the wireless network (110) using multiple digital twin instances representing the plurality of access points (111, 112, 113) and one or more other access points (114, 115). Components for determining one or more optimized channel configurations for the plurality of access points (111, 112, 113) based at least on the network data and the selected digital twin method; and A component for applying one of the one or more optimized channel configurations to the plurality of access points (111, 112, 113) within the wireless network (110).
13. The apparatus of claim 12, wherein the component for selecting the digital twin method is configured to: The first digital twin method (131) is selected based on determining that the age of the network data is less than or equal to a threshold, or The second digital twin method (132) is selected based on determining that the age of the network data is higher than the threshold.
14. The apparatus of claim 12, wherein the component for determining the one or more optimized channel configurations is configured to: further determine the one or more optimized channel configurations based on channel pollution information. The channel pollution information indicates: Channel interference from one or more adjacent channels on one or more channels used by the plurality of access points (111, 112, 113), and Channel occupancy, which is measured as signal power observed by at least one of the plurality of access points (111, 112, 113) on the one or more channels.
15. A computer program product comprising instructions that, when executed by a device (500), cause the device (500) to perform at least the following: Network data is collected from multiple access points (111, 112, 113) within a wireless network (110), wherein the network data includes at least one of the following: performance information of the multiple access points (111, 112, 113), spatial information indicating the location of the multiple access points (111, 112, 113), or channel information indicating one or more channels used by the multiple access points (111, 112, 113); Based on the age of the network data, a digital twin method is selected from at least two predefined digital twin methods (131, 132) used to simulate the wireless network (110). The at least two predefined digital twin methods (131, 132) mentioned above include at least: A first digital twin method (131) is used to simulate the wireless network (110) using digital twins representing the plurality of access points (111, 112, 113), and A second digital twin method (132) is used to simulate the wireless network (110) using multiple digital twin instances representing the plurality of access points (111, 112, 113) and one or more other access points (114, 115). Based at least on the network data and the selected digital twin method, determine one or more optimized channel configurations for the plurality of access points (111, 112, 113); and One of the one or more optimized channel configurations is applied to the plurality of access points (111, 112, 113) within the wireless network (110).