QoE-aware dynamic resource allocation

By expressing the Wi-Fi resource allocation problem as PO-MDP and using machine learning models to estimate QoE, the access category and service priority are dynamically adjusted, solving the problem that QoS cannot reflect QoE and achieving optimization and fairness of user experience quality in Wi-Fi networks.

CN121793155APending Publication Date: 2026-04-03HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510469549.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-01
Filing Date
2025-04-15
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing Wi-Fi network resource allocation methods are mainly based on Quality of Service (QoS) metrics, but these metrics fail to accurately reflect the Quality of User Experience (QoE), resulting in an inability to optimize the end-user service experience.

Method used

An adaptive system and approach are adopted to formulate the Wi-Fi resource allocation problem as a partially observable Markov decision process (PO-MDP). A machine learning model is used to estimate QoE, and access categories and service priorities are dynamically adjusted through a policy agent to maximize the overall system QoE and QoE fairness.

Benefits of technology

Dynamically optimizing Wi-Fi resource allocation without relying on actual application sessions improves the Quality of User Experience (QoE) and achieves fairness and efficient utilization of QoE.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121793155A_ABST
    Figure CN121793155A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to quality of experience (QoE) aware dynamic resource allocation. Systems and methods are provided for maximizing / optimizing QoE associated with an application. A network controller may receive telemetry data from an access point (AP). A resource manager operably connected to the network controller may estimate QoE values for individual application traffic flows of the one or more application traffic flows delivered through the AP based on the telemetry data. Access categories and traffic priorities may be computed from the estimated QoE values and jointly assigned to a particular application traffic flow.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Wi-Fi networks have rapidly become a vital component of today's ubiquitous connectivity, enabling a wide variety of applications and services. These networks efficiently connect mobile devices (such as smartphones, tablets, and laptops) to the outside world, supporting various services and enabling interactive experiences and seamless data exchange. One aspect of Wi-Fi's popularity is its ability to support a wide range of deployments, from small to large. In one example, a Wi-Fi network may include a small residential deployment, while in another, it may be deployed extensively in large spaces such as airports, shopping malls, campuses, and stadiums. These large deployments can serve hundreds or even thousands of mobile devices simultaneously running diverse applications. For example, a single access point (AP) in an airport can support users engaged in streaming media content, browsing the internet, and making video calls, requiring high throughput and low latency. Similarly, a stadium connects thousands of users to a centralized controller managing millions of traffic streams for live streaming, social media, and real-time updates. Attached Figure Description

[0002] This disclosure has been described in detail with reference to the following accompanying drawings, based on one or more different examples. These drawings are provided for illustrative purposes only and depict only typical, non-limiting aspects of such examples.

[0003] Figure 1 The illustration shows examples of network configuration in which a learning-based network controller can be implemented, based on some of the disclosed techniques.

[0004] Figure 2 The illustration shows a resource manager and its associated operations used to guide a network controller in implementing QoE-aware dynamic resource allocation, based on some examples of the disclosed technology.

[0005] Figure 3 These are example computational components that can be used to implement Quality of Experience (QoE) estimation, based on some examples of the disclosed techniques.

[0006] Figure 4 The illustration shows an example action space based on some instances of the disclosed technology, in which access categories and service priorities are assigned to service flows.

[0007] Figure 5 The illustration shows example algorithms for training QoE-aware policy agents, based on some examples of the disclosed techniques.

[0008] Figure 6These are example computational components that can be used to implement learning-based network control, based on some examples of the disclosed technology.

[0009] Figure 7 It is an example computing component that can be used to implement examples of the disclosed technology.

[0010] These accompanying drawings are not exhaustive and do not limit this disclosure to the precise form disclosed. Detailed Implementation

[0011] As noted above, Wi-Fi is an integral part of today's internet infrastructure and utilizes various methods of resource allocation in Wi-Fi networks. However, traditional methods of resource allocation in Wi-Fi are often based on Quality of Service (QoS) metrics, which do not necessarily accurately reflect the user's Quality of Experience (QoE). As used herein, resource allocation can refer to the assignment or allocation of subcarriers within the channel bandwidth, which are grouped into Resource Units (RUs). These RUs can be assigned to different client devices or sites, allowing the AP to serve different client devices / sites during uplink or downlink transmissions. As will be described in more detail below, client devices or sites wishing to transmit data can transmit according to their access class, where data belonging to the same access class can be transmitted using specific data packets (e.g., multi-user (MU) orthogonal frequency division multiple access (OFDMA) packets).

[0012] Unlike Quality of Service (QoS), Quality of Service (QoE) focuses on the end-user device experience and the parameters that affect that experience. QoE measures whether an end-user's experience with a service (e.g., web browsing, phone calls, television broadcasts) is positive or negative; in other words, it's the end-user's response to service performance. Furthermore, while QoE focuses on the overall service experience, QoS is a description or measure of the overall performance of a service, focusing on the network infrastructure and the operational parameters that affect its transmission and reception. Therefore, QoE can be seen as a measure of the overall quality of the service provided from the end-user's perspective, while QoS typically focuses on the media or network itself, rather than the end-user's perspective. For example, QoS parameters can include, but are not limited to, packet loss, bit rate, throughput, transmission latency, availability, and jitter.

[0013] It should also be noted that the IEEE 802.11 series of wireless local area network (WLAN) technology standards (also known as Wi-Fi) typically include QoS extensions that can manage service prioritization based on the type of data / service. For example, some QoS extensions of the 802.11 protocol can prioritize the transmission of voice packets and video packets. In particular, Wi-Fi Multimedia (WMM) (formerly known as Wireless Multimedia Extension (WME)) is a subset of the 802.11e Wireless LAN (WLAN) specification that enhances QoS on the network by prioritizing packets (services) according to four access classes (ACs). According to WMM, the access classes (arranged from highest to lowest priority) include:

[0014] 1) Voice: By assigning the highest priority to voice packets, WMM enables concurrent Voice over IP (VoIP) calls with minimal latency and the highest quality;

[0015] 2) Video: By placing video packets in the second layer, WMM prioritizes them above all other data services and enables the support of three to four standard definition television (SDTV) streams or one high definition television (HDTV) stream over WLAN;

[0016] 3) Best-effort: Best-effort data packets include those from legacy devices or applications or devices lacking QoS standards; and

[0017] 4) Background: Background priority includes file downloads, print jobs, and other services that will not be affected by the increased latency.

[0018] Each of the aforementioned WMM access categories represents a different WLAN transmit and / or receive (Tx / Rx) policy. The WMM also defines how Differentiated Service Code Point (DSCP) values ​​are mapped to these access categories. For example, when a traffic flow (a sequence of data packets or packets originating from / arriving at a source / destination) travels from a wired network to a wireless client, the WMM maps DSCP values ​​to certain ACs, causing packets with different DSCP values ​​to be routed to different transmit queues. For instance, on the uplink (UL) side, applications on the client device can set DSCP values ​​for their packets based on application specifications. Before transmitting a traffic flow (also called an application flow) from an application, the flow scheduler can use the DSCP value to determine the Service ID (TID) that can be assigned to that traffic flow. The flow scheduler uses the Service ID to map that data packet and other data packets constituting that traffic flow to a corresponding queue in an AC. Therefore, packets in different transmit queues can be transmitted according to the different WLAN transmit policies of the ACs.

[0019] To address the aforementioned gap between QoS and QoE, the disclosed examples of techniques provide adaptive systems and methods that formulate the Wi-Fi resource allocation problem as a partially observable Markov decision process (PO-MDP) to maximize overall system QoE and QoE fairness. The disclosed examples of techniques can estimate QoE without using any application or client data. Instead, they leverage temporal dependencies in network telemetry data to estimate QoE using a machine learning model tailored to application flows. Policies can control the assignment of access classes and service priorities to application flows according to the estimated QoE and can be dynamically adjusted to handle different application classes and variable network conditions. A policy agent can control this assignment of access classes and service priorities, wherein the policy agent can be trained in a simulated environment using the same knowledge base developed (collected) during the training of the QoE estimation model. In this way, running real clients, generating services through actual application sessions, and collecting both QoE and telemetry data can be avoided.

[0020] More specifically, the disclosed technology examples pertain to resource managers that can be integrated with existing network controllers, such as WLAN controllers. WLAN controllers typically include management interfaces for configuring Wi-Fi networks, pushing configurations to access points (APs), etc. The resource manager can determine the appropriate access class and service priority to be assigned to application flows. The resource manager can use a probabilistic prediction ML model to estimate the QoE of application flows based solely on telemetry data at the WLAN controller (which already receives telemetry data from network elements, such as APs, and is therefore a suitable network element for integration with the resource manager), although other network elements may also be utilized. The QoE metric or characteristic is inferred by an application-specific ML model using a Long Short-Term Memory (LSTM) neural network. Reinforcement learning (RL) algorithms (using a Dual Deep Q-Network (DDQN) RL method and a feedforward neural network) can then be used to determine the access class and service priority for each application flow using the estimated application flow QoE. The WLAN controller can control parameters related to the radio link / channel between client devices and APs. Application flows can be assigned to one of three Access Classes (best-effort; video; or voice), and within each access class, they can be assigned low or high priority (or one or more other priority groups as needed) to control packet outflow rate. Based on the various possible combinations of access class and service priority designated as the "action space," QoE-optimized access class and service priority assignments for application flows can be performed.

[0021] It should be noted that the terms “optimized,” “best,” etc., used in this document can be used to mean making or achieving the most efficient or perfect performance possible. However, as those skilled in the art who are reading this document will recognize, perfection is not always achievable. Therefore, these terms can also include making or achieving the best possible, most efficient, or most practical performance in a given situation, or making or achieving better performance than that achievable with other settings or parameters.

[0022] Before describing in detail examples of the disclosed systems and methods, it is helpful to describe example network installations that may implement these systems and methods in various applications. Figure 1 The illustration shows an example of a network configuration 100 that can be implemented for organizations such as businesses, educational institutions, government entities, healthcare facilities, or other organizations. Figure 1 The illustration shows an example of a configuration implemented by an organization with multiple users (or at least multiple client devices 110) and at least one physical or geographic site 102. Network configuration 100 may include site 102 communicating with network 120.

[0023] Site 102 may include a main network, such as an office network, a home network, or other network installation. The main network may be a private network, such as a network that includes security and access controls to restrict access to authorized users of the private network. Authorized users may include company employees, residents, and business customers at site 102.

[0024] exist Figure 1 In the example, site 102 includes a controller 104 that communicates with network 120. Controller 104 can provide communication between site 102 and network 120. In addition to controller 104, site 102 may have other communication points with network 120. Although a single controller 104 is illustrated, site 102 may include multiple controllers and / or multiple communication points with network 120. In some examples, controller 104 may communicate with network 120 via a router. In other examples, controller 104 provides router functionality for devices in site 102. In this specification, the term "tunnel" refers to the encapsulation mode for transmitting data between the AP and the controller.

[0025] Controller 104 may be operable to configure and manage network devices, such as those at site 102. Controller 104 may be operable to configure and / or manage switches, routers, access points, and / or client devices connected to the network. Controller 104 itself may be an access point (AP), or provide access point (AP) functionality.

[0026] Controller 104 can communicate with one or more switches 108 and / or wireless access points (APs) 106A-C. Switches 108 and wireless APs 106A-C provide network connectivity for various client devices 110A-J. Using the connection to switch 108 or AP 106A-C, client devices 110A-J can access network resources, including other devices on the (site 102) network and network 120.

[0027] Examples of client devices may include: desktop computers, laptop computers, servers, web servers, authentication servers, authentication-authorization-accounting (AAA) servers, domain name system (DNS) servers, dynamic host configuration protocol (DHCP) servers, internet protocol (IP) servers, virtual private network (VPN) servers, network policy servers, mainframes, tablet computers, e-readers, netbook computers, televisions and similar monitors (e.g., smart TVs), content receivers, set-top boxes, personal digital assistants (PDAs), mobile phones, smartphones, smart terminals, dumb terminals, virtual terminals, video game consoles, virtual assistants, Internet of Things (IoT) devices, and so on.

[0028] Within site 102, switch 108 is included as an example of a network access point established for client device 110I-J within site 102. Client device 110I-J can connect to switch 108 and, through switch 108, can access other devices within network configuration 100. Client device 110I-J can also access network 120 through switch 108. Client device 110I-J can communicate with switch 108 via wired or wireless connections. In the illustrated example, switch 108 communicates with controller 104 via wired or wireless connection 112E.

[0029] Wireless AP 106A-C is another example included for establishing a network access point for client device 110A-H at site 102. Each of the APs 106A-C can be a combination of hardware, software, and / or firmware configured to provide wireless network connectivity to wireless client device 110A-H. Figure 1 In the example, AP 106A-C can be managed and configured by controller 104. AP 106A-C communicates with controller 104 and network 120 via connection 112A-D, which can be a wired or wireless interface.

[0030] Network 120 may be a public or private network (such as the Internet) or other communication network to allow connectivity between various sites (such as site 102) and access to servers (such as server 130). Network 120 may include third-party telecommunications lines, such as telephone lines, broadcast coaxial cables, fiber optic cables, satellite communications, cellular communications, etc. Network 120 may include any number of intermediate network devices, such as switches, routers, gateways, servers, and / or controllers, which are not directly part of network configuration 100 but facilitate communication between the various parts of network configuration 100 and between network configuration 100 and other network-connected entities.

[0031] In the example, the aforementioned resource manager can be embodied in server 130 and integrated with controller 104 to optimally assign network resources to client devices to optimize or maximize overall QoE and QoE fairness. As noted above and described in more detail below, server 130 (resource manager) can: extract relevant data from telemetry data obtained from the AP; infer / estimate the unobservable QoE of the application from the relevant telemetry data; and determine the optimal access category and service priority for the application flow.

[0032] Figure 2 The illustration shows a resource management system 200 according to some examples of the disclosed technology, and associated operations for implementing a learning-based network controller. In some examples, the resource management system 200 may include one or more applications running on a client device or station 210. An access point (AP) (such as AP 206) may serve the client device 210. Based on telemetry data from AP 206, the resource manager 230 may estimate the QoE of application flows delivered through AP 206 based on previous or historical observations of the telemetry data. The resource manager 230 may also learn adaptive resource allocation policies, which may be used to generate or modify application flow configurations that may be pushed to AP 206 by means of a controller (e.g., WLAN controller 204) controlling operational aspects of AP 206. Ultimately, data can be sent to, received from, or transmitted over a data network (such as the Internet 250).

[0033] like Figure 2As illustrated, data can flow along AP 206 to client device 210, and from client device 210 to WLAN controller 204, and via data paths including links 207 and 219 to the Internet 250. The telemetry processor 230A of resource manager 230 can receive telemetry data from AP 206 via Simple Network Management Protocol (SNMP) detector 204B of WLAN controller 204. SNMP detector 204B can query devices (such as AP 206) using the SNMP protocol and acquire event data by acting as a trap daemon and monitoring SNMP traps and events associated with the raw telemetry data (in this example). The raw telemetry data can include user, system, and radio metrics or operational characteristics, such as radio link parameters between client device 210 and AP 206, application service-related metrics, wireless capabilities, underlying Wi-Fi configuration, and overall system statistics. The raw telemetry data can be passed from SNMP detector 204B to telemetry processor 230A. The telemetry processor 230A can extract relevant data from raw telemetry data and, if necessary, convert the data into the required format. It should be noted that the relevant data extracted from the raw telemetry data is often implementation-specific and can vary depending on factors such as the type of application or service involved, network characteristics, etc. Table 2, discussed in more detail below, provides example characteristics representing this relevant data.

[0034] To reiterate, this telemetry data is the only data that Resource Manager 230 has access to. Resource Manager 230 does not need or use data from client device 210 or applications running on it.

[0035] A predictive ML model can be trained for each application class, and this predictive ML model can be loaded into the memory (not shown) of the resource manager 230. It should be understood that the application class can refer to typical QoS classes of Layer 2 and Layer 3, such as web browsing / email application classes, video conferencing application classes, etc. Video streaming applications, high-definition video streaming ( Application classes, etc. Application classes are typically associated with necessary throughput parameters, such as 500Kbps to 1Mbps for web browsing / email, 2-5Mbps for HD video streaming, and so on. In some examples, the application classes of interest are video streaming, video conferencing, and file transfer applications, but any application class can be considered as needed. Therefore, predictive ML models can be generated, trained, and loaded into resource manager 230. These predictive ML models can be represented by separate instances of QoE estimator 230B. Thus, when a data packet arrives at WLAN controller 204 from AP 206, flow classifier 204A can identify the appropriate application class corresponding to that data packet. The identified application class can be provided to telemetry processor 230A. Processed telemetry data from telemetry processor 230A can be forwarded to the corresponding application class instance of QoE estimator 230B. Multiple related instances of the QoE estimator 230B can estimate the QoE of the application flow reflected by processed telemetry data and can forward the estimated QoE information or data to the policy agent 230C.

[0036] Policy agent 230C can calculate access classes and service priorities for application flows based on estimated QoE metrics (determined as described above) to maximize overall QoE and QoE fairness. In some examples, multi-agent RL can be used, where a specific policy agent or policy agent instance serves a specific application class, such as video streaming, video conferencing, or file transfer. Instances of policy agent 230C can modify application flow configurations associated with their respective application classes via AP configuration processor 204C, which then pushes these policies to the AP (e.g., AP 206).

[0037] The Data Plane Engine 204D can handle the routing of data packets, for example, through tables or other mechanisms that can be used to determine, for example, the destination path of an incoming data packet or the data path through the network.

[0038] It should be noted that the exchange of access category and service priority information between the policy agent 230C and the configuration processor 204C, the transmission of raw indicator (telemetry) data from the SNMP detector 204B to the telemetry processor 230A, and the transmission of application class information from the flow classifier 204A to the telemetry processor 230A can be performed using, for example, the Transmission Control Protocol (TCP) and the necessary TCP interface between them. The transmission of processed telemetry data from the telemetry processor 230A to the QoE estimator 230B, and the transmission of application-specific QoE from the QoE estimator 230B to the policy agent 230C, can be performed via a shared memory connection. As used herein, a shared memory connection can refer to a shared memory region (not shown) used as a channel through which the telemetry processor 230A, QoE estimator 230B, and policy agent 230C can communicate with each other. Data exchange between the AP 206 and the WLAN controller 204 can be performed via the OpenFlow interface.

[0039] Estimating QoE presents challenges. Telemetry data is heterogeneous and can vary significantly in scale and format. Furthermore, telemetry data can include instantaneous values, cumulative statistics, and derived features, each with potential outliers and distinct domains, ranges, and distributions. Additionally, the relationship between telemetry and QoE is non-linear (and therefore complex) and based on hidden or latent features. Finally, QoE metrics or properties are often influenced by the temporal context and dependencies of the telemetry data. That is, the applied QoE can depend on the history of telemetry. Therefore, the examples of techniques disclosed aim to accurately model the temporal and hidden relationships between telemetry and QoE metrics based on historical telemetry data as noted above.

[0040] Figure 3 The illustration shows an example computational component 300 that can be used to implement QoE estimation, based on some examples of the disclosed techniques. See now. Figure 3 The computing component 300 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. Figure 3 In the example implementation, computing component 300 includes a hardware processor 302 and a machine-readable storage medium 304. In some examples, computing component 300 may be an embodiment of resource manager 230 (or its components, such as QoE estimator 230B).

[0041] Hardware processor 302 may be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for retrieving and executing instructions stored in machine-readable storage medium 304. Hardware processor 302 may fetch, decode, and execute instructions (such as instructions 306-316) to control processes or operations for application stream-specific QoE estimation. As an alternative to or supplement to retrieving and executing instructions, hardware processor 302 may include one or more electronic circuits comprising functional electronic components for executing one or more instructions, such as field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), or other electronic circuits.

[0042] Machine-readable storage media (such as machine-readable storage media 304) can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. Therefore, machine-readable storage media 304 can be, for example, random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), storage devices, optical discs, etc. In some examples, machine-readable storage media 304 can be a non-transitory storage medium, where the term "non-transitory" excludes transient propagation signals. As described in detail below, machine-readable storage media 304 can be encoded with executable instructions, such as instructions 306-316.

[0043] The hardware processor 302 can execute instruction 306 to change the number of applications running simultaneously on the client device in order to force a poorer QoE.

[0044] In some examples of the disclosed techniques, telemetry data can be processed sequentially to capture long-term dependencies while maintaining a reasonably long context over extended time periods. Specifically, examples of the disclosed techniques use a Long Short-Term Memory (LSTM) network to model QoE estimation, which captures trends over time and maintains representational gradient values ​​as the context length increases. It should be noted that other ML model designs can be used to model QoE estimation, but the prediction accuracy may vary.

[0045] To train an application-specific QoE estimation model based on examples of the disclosed technology, an application in a Wi-Fi testbed can be executed. Telemetry data and accurate QoE from video streaming, video conferencing, and file transfers can be collected based on the application execution. That is, the LSTM model can receive AP telemetry data as input and output an estimated QoE for the applicable application class.

[0046] Telemetry data collection and QoE estimation can be performed under diverse network conditions, ranging from underloaded to overloaded, by systematically modifying the Wi-Fi configuration. Table 1 lists examples of possible Wi-Fi control parameters that can be used to generate the modified Wi-Fi configuration. Furthermore, the number of concurrent applications running can be varied from 1 to n, where the choice of n ensures that all concurrently running application instances congest the network, resulting in poor QoE. Application sessions can be distributed across multiple Wi-Fi client devices (e.g., four client devices).

[0047]

[0048] Table 1

[0049] Hardware processor 302 can execute instructions 308 to run the application on the client devices. The application can be executed for each application session, in this example for four client devices. The duration of the application session can vary, but in one example, the application session can be set to 200 seconds. The application can be executed repeatedly as needed (e.g., twice in some examples) to increase the size and diversity of the dataset.

[0050] Hardware processor 302 can execute instructions 310 to collect telemetry data associated with application execution and measure the QoE of the application at the client device. Using the application execution method described above, a large amount of telemetry data can be collected, for example, covering applications running for a total of 251 hours across all client devices for three application classes (video streaming, video conferencing, and file transfer applications), resulting in over 226,000 data points. Telemetry data can be collected according to some time specification (e.g., periodic collection, once every four seconds).

[0051] Furthermore, QoE can be measured during application execution on the client device. The application's QoE metric can be calculated at each interval, and this metric can be normalized for the maximum achievable QoE value for each application. In this way, QoE values ​​can be compared across application classes. In particular, while subjective metrics are often the most accurate indicators of user experience, collecting subjective metrics is impractical for a large number of system configurations. Therefore, objective quality metrics can be used, which reasonably approximate the end-user's perceived quality.

[0052] To calculate the QoE of a video stream, an algorithm based on Model Predictive Control (MPC) can be used, where the QoE metric is a function of the following: the quality of each video chunk, the quality variation between consecutive chunks, rebuffering time, and startup latency. The principle behind this approach is that once video playback begins, users tend to prefer uninterrupted playback (without pauses) and do not want frequent quality variations between video chunks. Using MPC, the optimal possible QoE is achieved throughout the entire video streaming session.

[0053] To measure QoE in video conferencing, the following metrics are key considerations: audio jitter; audio packet loss; video packet loss; and transmitted video packets. The rationale for considering these metrics is that end users generally find choppy, out-of-sync live audio or video frames unpleasant and undesirable. The best possible QoE for a video conferencing application is achieved using the highest possible video resolution, the highest possible frame rate, and the lowest possible jitter and packet loss. For video conferencing applications, QoE values ​​are normalized against this best possible QoE score. To map these QoE metrics to a four-second interval, QoE values ​​can be averaged over four seconds, and individual QoE values ​​can be reported.

[0054] Regarding file transfer QoE, since QoE is calculated at each time interval (as described above) (rather than simply using the file transfer completion time), the number of bytes transferred from the server to the client within one second can be recorded. This data can then be averaged over four-second intervals to calculate the QoE for that time period, aligning it with a specified four-second telemetry collection interval. The rationale for calculating file transfer QoE in this way is that end users tend to want file transfers to complete in the shortest possible time, which is achievable if the server is continuously / constantly sending large amounts of data to the receiver.

[0055] Hardware processor 302 can execute instructions 314 to align telemetry data associated with application execution with application-stream-specific QoE measurements. In some examples, QoE and telemetry data are synchronized based on their timestamps. Furthermore, since each application can report its QoE at a different rate than the telemetry, these measurements can be aligned by averaging the QoE metrics over every four-second interval. By mapping the average QoE metric from t-4 to t seconds to the corresponding telemetry, this ensures that the telemetry data at time t accurately reflects the system state.

[0056] Furthermore, alignment telemetry data can include data normalization. In the case of video streaming applications, the QoE of each (as described above) video chunk can be normalized for the best possible QoE score. QoE values ​​range from negative infinity to 1, although experimentally, QoE typically does not fall below negative 1. Therefore, the range of QoE values ​​can be limited to between negative 1 and 1. A min-max normalization method can be used to normalize these QoE values ​​between 0 and 1. This normalized QoE is reported for each video chunk, corresponding to a four-second interval. It should be noted that video streaming applications operate in a balanced mode to maintain a balance between avoiding rebuffering (pauses) and avoiding instability (video quality variations). This achieves minimal rebuffering and stable quality.

[0057] For video conferencing applications, the best possible QoE is typically achieved using the highest possible video resolution, the highest possible frame rate, and the least jitter and packet loss. For video conferencing applications, the QoE value is normalized to this best possible QoE score.

[0058] For file transfer applications, the optimal achievable QoE for file transfer can be determined using an optimal Wi-Fi configuration, which includes a 160MHz channel using the 5GHz band and an HE physical interface. Under this setting, the maximum number of bytes the aforementioned server can transfer to the client device represents the optimal achievable QoE. The QoE at four-second intervals can be normalized to this maximum file transfer QoE score.

[0059] Hardware processor 302 can execute instructions 314 to train a QoE model. Recall that QoE ML models can be developed and implemented for each application class. To train an application-specific QoE model, an LSTM network can be constructed using, for example, an input layer with 257 telemetry data items and three stacked LSTM layers, where the first, second, and third layers have 256, 128, and 64 hidden units, respectively. Dense output layers can follow the LSTM layers to compute the estimated QoE metric. The application-specific QoE model can be trained using the Adam optimizer and the mean squared error (MSE) loss function. The dataset can be split into training, validation, and test datasets using split ratios of 60%, 20%, and 20%, respectively. The training process can include 50 epochs with a batch size of 32. Table 2 (see below) highlights example features identified by the model using a permutation feature importance method for QoE estimation across all applications.

[0060]

[0061] Table 2

[0062] Although the QoE estimator 230B is used to estimate the state of the PO-MDP, the transition probabilities remain difficult to model due to the complex relationship between the actions taken and the observed states. Examples of the disclosed techniques address the complexity of using RL in large search spaces and the potential lack of sufficient training data by employing RL methods that explore various possibilities within the search space while effectively utilizing the collected states. The exploration process can be controlled to accelerate model convergence, while DNNs handle the high-dimensional state-action space in Wi-Fi settings, caused by varying application classes, dynamic network conditions, and diverse client capabilities. As noted above, the policy agent 230C can leverage a DDQ NRL method combined with a feedforward neural network to handle the high-dimensional space and uses a replay buffer to improve sample efficiency.

[0063] More specifically, the Q-network, denoted by Q(s,a;θ), is a neural network that approximates the action-value function, where s represents the state, a represents the action, and θ represents the network parameters. The target Q-network Q′ is a copy of the Q-network, which is periodically updated to stabilize the learning process and minimize the potential overestimation of the Q-value. During training, the Q-value is updated as follows:

[0064] Q(s t ,a t )←Q(s t ,a t )+α[r t +γQ′(s t+1 argmax a′ Q(s t+1 ,a′;θ)θ′)-

[0065] Q(s t ,a t ;θ)],(Equation 1)

[0066] Where s t It is the current state at time t, a t The action taken at time t, a′ is the action that maximizes the Q value of the next state, and r t The reward received at time t, s t +1 is the next state at time t+1, θ′ is the target Q-network parameter, α is the learning rate, and γ is the discount factor. Boltzmann exploration can be used to select actions, assigning a probability to each action based on its Q-value, as follows:

[0067]

[0068] Where T is the temperature parameter for controlling the exploration-utilization tradeoff.

[0069] The loss function used to train the Q-network is the mean squared error (MSE) between the predicted Q-value and the target Q-value:

[0070] L(θ)=E[(r t +γQ′(s t+1 argmax a′ Q(s t+1 ,a′;θ)θ′)-Q(s t ,a t ;θ)) 2 ].

[0071] (Equation 3)

[0072] As noted above, the disclosed technology example controls parameters related to the radio link between Wi-Fi client devices and the access point (AP). Conventional WLAN controllers configure flow-level, radio-level, and device-level parameters. Since configuring radio-level and device-level parameters is associated with restarting the AP or device, the resource manager 200 operates to configure flow-level parameters that can be modified at runtime. Runtime modifications according to the disclosed technology example provide the ability to achieve desired improvements in QoE for flows while maintaining, for example, QoE fairness across multiple flows. These parameters can be calculated by the resource manager 200 and ultimately modified via the configuration processor 204C of the WLAN controller 204.

[0073] Access class and service priority can be assigned to application flows jointly or in combination. When assigning access class aspects to application flows, the application class of the application flow is not limited to a specific access class. Instead, policy agent 230C can assign incoming application flows to any access class based on learned policies to adapt to application heterogeneity and diverse network conditions. As mentioned above, in some examples, application flows can be assigned to any of the three access classes: Best Effort (BE), Video (VI), and Voice (VO). Furthermore, policy agent 230C can assign low priority (LQ) or high priority (HQ) to each flow, which controls the packet ejection rate. These actions (i.e., assigning AC / service priority combinations to each application) can be performed at a given period, such as once every 16 seconds. This periodic execution of actions allows policy agent 230C (and resource manager 230 as a whole) to dynamically adapt to changing network conditions. The duration / period of execution can, for example, depend on the processor of WLAN controller 204 ( Figure 2 The processing capacity and speed (not shown in the figure) vary.

[0074] Figure 4The illustration shows an example action space 400, which specifies the access class / service priority combinations that can be assigned to incoming application flows. Here, "action" as used herein can refer to the action of assigning possible access class / service priority combinations to application flows. This is achieved by connecting low-priority and high-priority queues (LQ / HQ) with... Figure 4 The access classes shown, combined with examples of the disclosed technology, allow action space 400 to include six possible action combinations: LQ+BE and HQ+BE for BE access class 402; LQ+VI and HQ+VI for VI access class 404; and LQ+VO and HQ+VO for VO access class 406. A policy agent (such as policy agent 230C) can assign specific access class / service priority combinations defined in action space 400 to incoming data packets representing one or more application flows 408. As described above, the application flow configuration is changed / updated via configuration processor 204C by periodically (e.g., every 16 seconds) assigning access class / service priorities to the application flows using the policy agent. The configuration processor 204C can then push these policies to the APs (e.g., AP 206) that the application flows are traversing.

[0075] It should be noted that the smallest unit of action taken by the policy agent (called the state space) is a 5-tuple application flow traversing the AP. Given the complexity and volume of the available data in the controller database (containing thousands of metrics), identifying which of these metrics have a significant impact on QoE estimation can be challenging. To ensure comprehensive coverage, all available metrics are considered based on the examples of the disclosed techniques, which allow the QoE estimator to select the relevant subset and accurately estimate the QoE.

[0076] AP telemetry data can be categorized into one of three data types: application-specific data, client-specific data, and access point-specific data. Application-specific data can include metrics such as stream-specific throughput, stream transport protocol, and TX / RX bytes. Client-specific data can contain client radio capabilities, such as signal-to-noise ratio (SNR), PHY / MAC interface details, channel width, and access class. Access point-specific data can relate to system-specific metrics, such as total packet loss, packet drop, buffer TX / RX metrics, and various radio metrics. The metrics mentioned above are merely examples to provide context, as many other related metrics can be considered within each data category.

[0077] As noted above, the policy agent configured according to the examples of the disclosed technology utilizes a DDQN RL method coupled with a feedforward neural network. In some examples, the reward function of DDQN is configured to maximize the overall QoE and maximize QoE fairness. Due to the use of multi-agent RL mentioned earlier (recall that different policy agent instances serve different application classes), each policy agent instance receives, in addition to the state: (1) the average QoE of all other applications; (2) the deviation of the QoE of the specific application from the average QoE. In state s t Take action a t Then, the reward function can be defined as:

[0078]

[0079] QoE app It is the QoE of the current application stream. It is the average QoE of all other application streams, and It is the deviation between the current application's QoE and the system-wide average QoE, where This is the system-wide average QoE for N application flows. Weight w QoE w mean and w deviation It is used to adjust the application's QoE sensitivity to the Wi-Fi environment, thus allowing the reward function to be customized according to the specific needs of different applications. This function maximizes the QoE of the current application, keeping it in balance with the average QoE of all other applications, and maximizes QoE fairness by minimizing the bias in the QoE metric.

[0080] Further concerning the policy agent in Resource Manager, training the policy agent in Maestro presents several challenges. The training process involves observing state variables, executing actions, and accessing rewards and the next state. Implementing this in a real / effective Wi-Fi setup or environment requires running real clients, generating services through actual application sessions, and collecting QoE and telemetry data. In other words, due to the complexity and variability of real-world network conditions, obtaining a sufficient number of diverse samples for the policy agent to learn access category / service priority assignment policies is impractical. Furthermore, these real-world Wi-Fi conditions cannot be easily integrated with current RL environments, not to mention exploring all possible state-action combinations is infeasible due to the significant time and resources required.

[0081] Therefore, based on the examples of the disclosed technology, a simulation environment can be used to train the policy agent, which utilizes an extensive knowledge base already collected during the training of the QoE estimator. This knowledge base includes telemetry data and corresponding QoE values, and enables the resource manager to effectively simulate the training process of the policy agent.

[0082] Figure 5 The illustration shows an example algorithm 500 representing an iterative learning process employed by a policy agent in a simulated environment. The policy agent begins with a random state within a knowledge base. Once an action is selected based on this state, algorithm 500 evaluates the action by comparing it to the knowledge base. Algorithm 500 uses the Pearson correlation coefficient ρ to find the optimal matching vector v. That is, by maximizing ρ, algorithm 500 identifies the optimal matching vector v corresponding to the current state, Wi-Fi configuration, the action taken, and the reward received. This matching vector v represents the next state, allowing the policy agent to iteratively continue its learning process.

[0083] To estimate the weights of the reward function in Equation 4 and the temperature values ​​of the Boltzmann exploration in Equation 2, a representational set of five different reward function weights and ten temperature values ​​(ranging from 0 to 1, incrementing by 0.1) can be used. These combinations of values ​​can be input into Algorithm 500 one at a time. The final weights and temperature values ​​based on the loss function can then be selected.

[0084] It should be noted that the loss function for RL training follows a behavior similar to the function L(t) = α*t*e⁻βt, where t is the training time or number of epochs, α is a scaling factor determining the height of the peak, and β is a factor controlling the rate of decline after the peak. This is because the function represents the expected loss value over time. This loss value initially rises as the policy agent explores, reaches a peak, and then gradually decreases to zero as the policy agent learns and optimizes its (access category / business priority assignment) policy. The reward weights and temperature values ​​can be selected by fitting this function to the observed loss value.

[0085] Based on a state size of 257, an action (space) size of 6, a hidden layer size of 128, a discount factor of 0.99, a batch size of 32, an Adam optimizer, a learning rate of 0.0001, and a replay buffer of 10,000, Table 4 lists the optimization reward weights and temperature values ​​for a set of examples for each application class.

[0086] application <![CDATA[w QoE ]]> <![CDATA[w mean ]]> <![CDATA[w deviation ]]> T video stream 0.2 0.6 0.2 0.4 videoconference 0.2 0.2 0.6 0.6 File transfer 0.34 0.33 0.33 0.8

[0087] Table 4

[0088] Figure 6 The illustration shows an example computing component 600 that can be used to implement QoE-aware dynamic resource allocation, based on some examples of the disclosed technology. See now. Figure 6 Computing component 600 can be, for example (similar to computing component 300) a server computer, controller, or any other similar computing component capable of processing data. Figure 6 In example implementations, like computing component 300, computing component 600 includes a similar hardware processor 602 and a similar machine-readable storage medium 604. In some examples, computing component 600 may be an embodiment of resource manager 230, server (such as server 160), etc.

[0089] As described in detail below, the machine-readable storage medium 604 may be encoded with executable instructions, such as instructions 606-612.

[0090] Hardware processor 602 can execute instructions 606 to receive telemetry data from the AP at the network controller. As previously discussed, client devices can be served by the AP while running applications such as video streaming applications, file transfer applications, etc. Unlike other systems and methods that require collection (e.g., internal application status information or other data) from client devices(s) or applications(s), the disclosed technology example is able to estimate QoE individually based on telemetry data associated with application operation on client devices whose data packets / traffic pass through the AP. As discussed above, telemetry data can be parsed to extract relevant data, and telemetry data can be normalized / transformed to account for variations between types.

[0091] Hardware processor 602 can execute instructions 608 to estimate the QoE value for individual application traffic flows in one or more application traffic flows delivered through the AP, based on telemetry data. The telemetry processor, which incorporates telemetry data from the AP and performs the aforementioned relevant data extraction and normalization, can pass the relevant data to a QoE estimator. The QoE estimator can include multiple QoE estimator instances used to estimate the QoE of applications running on client devices based on the application class(s) to which the applications(s) are categorized, as demonstrated by the received telemetry data. The QoE estimator instances include ML models trained to accurately model the temporal and hidden relationships between telemetry data and QoE metrics. That is, the QoE estimator instances infer "unobservable" QoE metrics. In some examples, LSTM networks can be constructed to train these QoE ML models, where training data is derived from telemetry and measured QoE information under diverse network conditions, and where the telemetry data and measured QoE information are aligned, for example, based on timestamps.

[0092] Hardware processor 602 can calculate access class and service priority based on estimated QoE values ​​for individual application service flows, and assign access class and service priority to individual application service flows. In some examples, multiple policy agents (which include ML models also for specific application classes) are trained to maximize overall QoE and QoE fairness within the network and AP / client devices. Reinforcement learning can be used to perform this assignment. Finally, the policy agent for an application class assigns actions to application flows belonging to its application class. This action includes assigning both access class and service priority to the application flow. It should be noted that application flows can be assigned to any access class based on policies learned by policy agent instances. It should also be noted that policy agent instances can be trained in a simulated environment to avoid the operational costs associated with running actual client devices and applications and changes in Wi-Fi network operations. The simulated environment can utilize an existing knowledge base that includes received telemetry data and corresponding QoE values.

[0093] Hardware processor 602 can execute instructions 612 to maximize the overall QoE across one or more application flows and the QoE parity applied to individual application flows by adjusting the configuration of individual application flows according to their respective assigned access categories and service priorities. As noted above, in some examples, configuration actions can occur / be executed at given intervals to account for dynamic network conditions. The configuration handler operates to push the policies / configurations determined / assigned by the policy broker instance to the AP.

[0094] Figure 7 A block diagram of an example computer system 700 is depicted in which various examples of the disclosed technologies described herein may be implemented. The computer system 700 includes a bus 702 or other communication mechanism for transmitting information, and one or more hardware processors 704 coupled to the bus 702 for processing information. For example, the hardware processors 704 may be one or more general-purpose microprocessors.

[0095] Computer system 700 may include main memory 706, such as random access memory (RAM), cache, and / or other dynamic storage devices, coupled to bus 702, for storing information and instructions to be executed by processor 704. Main memory 706 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 704. When stored in a storage medium accessible to processor 704, such instructions turn computer system 700 into a special-purpose machine customized to perform the operations specified in the instructions.

[0096] The computer system 700 also includes a read-only memory (ROM) 708 or other static storage device coupled to the bus 702 for storing static information and instructions of the processor 704. Storage devices 710 (such as disks, optical discs, or USB thumb drives (flash drives)) are provided and coupled to the bus 702 for storing information and instructions.

[0097] Computer system 700 can be coupled to display 712, such as a liquid crystal display (LCD) (or touchscreen), via bus 702 for displaying information to the computer user. Input device 714 (including alphanumeric and other keys) is coupled to bus 702 for transmitting information and command selections to processor 704. Another type of user input device is cursor control 716 (such as a mouse, trackball, or arrow keys) for transmitting directional information and command selections to processor 704 and for controlling cursor movement on display 712. In some examples, the same directional information and command selections as with cursor control can be achieved by receiving touches on a touchscreen without a cursor.

[0098] The computing system 700 may include a user interface module to implement a GUI, which may be stored as executable software code executed by (or more) computing devices in a mass storage device. For example, this module and other modules may include components such as software components, object-oriented software components, class components and task components, processes, functions, properties, programs, subroutines, program code segments, drivers, firmware, microcode, circuit systems, data, databases, data structures, tables, arrays, and variables.

[0099] Computer system 700 can embody components including a QoE-aware dynamic resource allocation system, for example, Figure 2 Components such as file explorer and network coordinator.

[0100] As used herein, the term "non-transitory medium" and similar terms refer to any medium that stores data and / or instructions that cause a machine to operate in a specific manner. Such non-transitory media can include non-volatile media and / or volatile media. For example, non-volatile media include optical discs or magnetic disks such as storage device 710. For example, volatile media include dynamic memory such as main memory 706. Non-transitory media are distinct from transmission media, but can be used in conjunction with transmission media. Transmission media participate in the transfer of information between non-transitory media.

[0101] Computer system 700 also includes a communication interface 718 coupled to bus 702. Network interface 718 provides bidirectional data communication coupled to one or more network links connected to one or more local networks. For example, communication interface 718 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem to provide data communication connectivity with a corresponding type of telephone line. A wireless link may also be implemented. In any such implementation, network interface 718 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.

[0102] Each process, method, and algorithm described in the preceding sections can be embodied in code components executed by one or more computer systems or computer processors, including computer hardware, and can be fully or partially automated by these code components. One or more computer systems or computer processors can also operate to support the execution of related operations in a “cloud computing” environment or to perform related operations as “Software as a Service” (SaaS). Processes and algorithms can be implemented, partially or entirely, in application-specific circuit systems. The various features and processes described above can be used independently of each other or combined in various ways. Different combinations and sub-combinations are intended to fall within the scope of this disclosure, and certain method or process boxes may be omitted in some implementations. The methods and processes described herein are not limited to any particular order, and the boxes or states associated with them can be executed in other suitable orders, or in parallel, or in some other way. Boxes or states can be added or removed from the examples disclosed. The execution of certain operations or processes can be distributed among computer systems or computer processors, not residing within a single machine, but deployed across several machines.

[0103] As used herein, the term “or” can be interpreted as inclusive or exclusive. Furthermore, singular descriptions of resources, operations, or structures should not be interpreted as excluding the plural. Unless explicitly stated otherwise or understood otherwise in the context in which they are used, conditional language (such as “can,” “will,” “may,” or “may”) is generally intended to convey that some examples include certain features, elements, and / or steps, while other examples do not.

[0104] Unless otherwise expressly stated, the terms and phrases used in this document, and their variations thereof, should be interpreted as open-ended rather than restrictive. Adjectives and similar terms such as “regular,” “traditional,” “normal,” “standard,” “known,” and similar meanings should not be interpreted as limiting the item to a given time period or to items available at a given time, but should be interpreted as encompassing regular, traditional, normal, or standard techniques available or known at any time now or in the future. In some cases, the presence of expansive words and phrases such as “one or more,” “at least,” “but not limited to,” or other similar phrases should not be interpreted as implying an intention to or requirement of a narrower definition where such expansive words might not be present.

Claims

1. A method comprising: Receive telemetry data from the access point (AP) from the network controller of the network; The Quality of Experience (QoE) value is estimated based on the telemetry data for individual application service flows in one or more application service flows delivered through the AP. The access category and service priority are calculated based on the estimated QoE value, and the access category and service priority are assigned to the individual application service flow. By adjusting the configuration of the individual application service flow according to its assigned access category and service priority, the overall QoE across the one or more application service flows and the QoE balance applied to the individual application service flow are maximized.

2. The method of claim 1, wherein the received telemetry data includes raw telemetry data, the raw telemetry data including user metrics, network metrics of the network, and radio metrics of one or more radios operating in the AP.

3. The method according to claim 2, further comprising: The raw telemetry data is processed to extract telemetry data related to QoE estimation.

4. The method according to claim 3, wherein the processing of the raw telemetry data includes: Process sequential telemetry data.

5. The method of claim 1, wherein the estimation of the QoE value comprises: Identifies the application class associated with the packet of the one or more application service flows received by the network controller.

6. The method according to claim 5, further comprising: The identified application class is used to apply an application class-specific prediction model corresponding to the identified application class to estimate the QoE value based on the telemetry data.

7. The method of claim 6, wherein the application-specific prediction model comprises a Long Short-Term Memory (LSTM) neural network.

8. The method according to claim 6, further comprising: The application-class-specific prediction model is trained using telemetry data and measured QoE collected under various network conditions, including underloaded and overloaded network conditions.

9. The method according to claim 8, further comprising: The collected telemetry data and the measured QoE are synchronized based on the corresponding timestamps associated with the collected telemetry data and the measured QoE.

10. The method of claim 1, wherein calculating the access class and service priority based on the estimated QoE value and assigning the access class and service priority to the individual application service flow is performed by an application-class-specific policy agent.

11. The method of claim 1, wherein the application-class specific policy agent uses a dual deep Q-network DDQN reinforcement learning RL algorithm coupled to a feedforward neural network.

12. The method of claim 11, wherein calculating the access class and service priority based on the estimated QoE value and assigning the access class and service priority to the individual application service flow comprises: The access category and service priority are jointly assigned based on the action space, which includes possible combinations of access categories and service priorities.

13. The method of claim 10, further comprising: Train the application-specific policy agent in a simulated environment.

14. A system comprising: processor; as well as A memory unit, the memory unit including instructions, which, when executed, cause the processor to: Receive telemetry data from the access point (AP); Based on the telemetry data, estimate the Quality of Experience (QoE) value for each of the multiple application service flows passing through the AP; Based on the estimated QoE value, the access category and service priority are calculated for each application service flow; The calculated access category and service priority are jointly assigned to each application flow; as well as The configuration, including the jointly assigned access category and service priority, is pushed to the network controller to be forwarded to the AP for reconfiguration, thereby processing subsequent data packets belonging to each application service flow in the application service flow according to the jointly assigned access category and service priority.

15. The system of claim 14, wherein the instruction that causes the processor to estimate the QoE value when executed includes instructions that, when executed, also cause the processor to perform the following operations: Identify the application class associated with the data packets of the one or more application service flows received by the network controller; and An application-class-specific prediction model corresponding to the identified application class is applied to estimate the QoE value based on the telemetry data.

16. The system of claim 15, wherein the application-specific prediction model comprises a Long Short-Term Memory (LSTM) neural network.

17. The system of claim 14, wherein the memory unit further includes instructions that, when executed, also cause the processor to perform the following operations: train the application-specific prediction model using telemetry data collected and measured QoE under various network conditions including underload and overload network conditions.

18. The system of claim 14, wherein the instructions that, when executed, cause the processor to calculate and jointly assign the access category and service priority further include, when executed, instructions that also cause the processor to perform the following operation: execute a policy proxy instance specific to each identified application class.

19. The system of claim 18, wherein the policy agent instance agent uses a dual deep Q network (DDQN) reinforcement learning RL algorithm coupled with a feedforward neural network to calculate the access category and service priority.

20. A system comprising: processor; as well as A storage unit, the storage unit including instructions, which, when executed, cause the processor to: The number of applications to be run simultaneously on a client device is changed to force a poorer Quality of Experience (QoE) that will be experienced by the client device, which is served by an access point (AP) that is operational in the network. The application is executed on the client device; Telemetry data is collected from the AP, and QoE is measured, the telemetry data being associated with the execution of the application; as well as The telemetry data and the measured QoE are used to train a QoE model to be operationalized to maximize the overall QoE across multiple traffic flows traversing the AP and the QoE balance of individual traffic flows applied to the multiple traffic flows, each of the multiple traffic flows being associated with one of multiple applications executed at the client device.