QOE-AWARE DYNAMIC RESOURCE ALLOCATION

By integrating a resource manager with machine learning and reinforcement learning to estimate QoE from telemetry data, the method optimizes Wi-Fi resource allocation, addressing the gap between QoS and QoE, and enhancing user satisfaction and network performance.

DE102025111138A1Pending Publication Date: 2026-04-02HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-22
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Conventional Wi-Fi resource allocation methods rely on Quality of Service (QoS) metrics, which do not accurately reflect the end-user's Quality of Experience (QoE), leading to suboptimal network performance and user satisfaction.

Method used

A resource manager integrated into a WLAN controller estimates QoE using machine learning models based on network telemetry data, employing a probabilistic predictive model and reinforcement learning to dynamically allocate access categories and traffic priorities, optimizing QoE and fairness across various applications and network conditions.

Benefits of technology

The solution effectively maximizes overall system QoE and fairness by accurately estimating user experience without requiring real application sessions, adapting to dynamic network conditions, and ensuring seamless data exchange for multiple users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems and procedures are provided to maximize / optimize the quality of experience (QoE) associated with applications. A network controller can receive telemetry data from an access point (AP). A resource manager, operationally connected to the network controller, can estimate a QoE value for individual application traffic streams from one or more application traffic streams passing through the AP, based on the telemetry data. An access category and traffic priority, based on the estimated QoE value, can be calculated and jointly assigned to a specific application traffic stream.
Need to check novelty before this filing date? Find Prior Art

Description

background

[0001] Wi-Fi networks have rapidly become a crucial element of today's ubiquitous connectivity, enabling a multitude of applications and services. These networks efficiently connect mobile devices such as smartphones, tablets, and laptops to the outside world, supporting a wide range of services and facilitating interactive experiences and seamless data exchange. One aspect of Wi-Fi's popularity is its ability to support a broad spectrum of deployments, from small to large. For example, a Wi-Fi network can be set up in a small apartment building, while in other scenarios, it can be deployed in large spaces such as airports, shopping malls, universities, and stadiums. These large installations can serve hundreds or even thousands of mobile devices running various applications simultaneously.A single access point (AP) in an airport, for example, can support users streaming media, browsing the internet, and making video calls, requiring both high throughput and low latency. Similarly, in stadiums, thousands of users are connected to a central controller that manages millions of data streams for live streaming, social media, and real-time updates. Brief description of the drawings

[0002] The present disclosure is described in detail in accordance with one or more different examples, with reference to the following illustrations. The illustrations serve only for clarification and merely represent typical, non-limiting aspects of such examples. Fig.shows an example of a network configuration in which a learning network controller can be implemented in accordance with some examples of the disclosed technology. Fig. shows a resource manager and related operations for controlling a network controller to achieve QoE-aware dynamic resource allocation in accordance with some examples of the disclosed technology. Fig. is an example of a computer component that can be used to implement the estimation of Quality of Experience (QoE) in accordance with some examples of the disclosed technology. Fig. shows an example of an action space in which access categories and traffic priorities are assigned to traffic flows in accordance with some examples of the disclosed technology. Fig.shows an example algorithm for training a QoE-enabled policy agent in accordance with some examples of the disclosed technology. Fig. is an example of a computer component that can be used to effect learning-based network control in accordance with some examples of the disclosed technology. Fig. is an example of a computer component that can be used to implement examples of the disclosed technology.

[0003] The illustrations are not exhaustive and do not limit the present disclosure to the exact form that is disclosed. Detailed description

[0004] As previously mentioned, Wi-Fi is an integral part of today's internet infrastructure, and various approaches are used for resource allocation in a Wi-Fi network. However, conventional approaches to resource allocation in Wi-Fi are often based on Quality of Service (QoS) metrics, which do not necessarily accurately reflect a user's Quality of Experience (QoE). The resource allocation used here can refer to the allocation of subcarriers within a channel bandwidth, grouped into Resource Units (RUs). Such RUs can be assigned to different client devices or stations, allowing access points (APs) to serve the various client devices / stations during uplink or downlink transmissions.As described in more detail below, client devices or stations that wish to send data can do so in accordance with access categories, whereby data belonging to the same access category can be transmitted using specific data packets, e.g., Multi-User (MU) Orthogonal Frequency Division Multiple Access (OFDMA) packets.

[0005] Quality of Experience (QoE), which differs from Quality of Service (QoS), focuses on the experience on end-user devices and the parameters that affect that experience. QoE is a measure of an end-user's satisfaction or dissatisfaction with a service (e.g., web browsing, phone calls, television broadcasts); in other words, an end-user's response to the service's performance. While QoE focuses on the overall service experience, QoS is a description or measurement of a service's overall performance, focusing on the network infrastructure and the operational parameters that affect transmission / reception within that infrastructure. Thus, QoE can be a measure of the overall quality of the service offered from the end-user's perspective, whereas QoS generally focuses on the media or the network itself—not the end-user's perspective.QoS parameters include, but are not limited to, packet loss, bit rate, throughput, transmission delay, availability, jitter, etc.

[0006] It should also be noted that the IEEE 802.11 family of standards for wireless local area networks (WLANs), also known as Wi-Fi, typically includes QoS extensions that can manage traffic prioritization based on the type of data. For example, QoS extensions for some 802.11 protocols can prioritize the transmission of voice and video packets. Specifically, Wi-Fi Multimedia (WMM), formerly known as Wireless Multimedia Extensions (WME), is a subset of the 802.11e specification for wireless LANs (WLANs) that enhances QoS in a network by prioritizing packets (traffic) according to four Access Categories (ACs). According to WMM, these Access Categories (from highest to lowest priority) include: 1) Voice: By giving highest priority to voice packets, WMM enables simultaneous Voice over IP (VoIP) calls with minimal latency and the best possible quality; 2) Video: By classifying video packets into the second layer, WMM prioritizes them over other data traffic and enables support for three to four standard-definition TV streams (SDTV) or one high-definition TV stream (HDTV) in a WLAN; 3) Best effort: Best-effort data packets consist of packets originating from older devices or from applications or devices that do not have QoS standards; and 4) Background: Background priority includes file downloads, print jobs, and other data traffic that does not suffer from increased latency.

[0007] Each of the aforementioned WMM access categories represents a different WLAN transmit and / or receive (Tx / Rx) policy. WMM also defines how Differentiated Services Code Point (DSCP) values ​​can be categorized into these access categories. For example, when traffic flows (contiguous data packets or a sequence of data packets going to / from a source / destination) travel from a wired network to a wireless client, WMM assigns DSCP values ​​to specific access points (ACs) so that packets containing different DSCP values ​​are routed to different transmission queues. On the uplink (UL) side, for instance, an application on a client device can specify a DSCP value for its packets based on the application's specifications.Before transmitting a traffic flow from the application (also known as an application flow), a flow planner can use the DSCP value to determine a traffic ID (TID) that can be assigned to the traffic flow. The flow planner uses this TID to assign the data packet and other data packets that make up the traffic flow to a queue corresponding to one of the ACs. In this way, the packets in the various transmission queues can be transmitted in accordance with the different WLAN transmission policies of the ACs.

[0008] To bridge the aforementioned gap between QoS and QoE, disclosed technology examples offer adaptive systems and methods that frame the Wi-Fi resource allocation problem as a partially observable Markov decision process (PO-MDP) to maximize overall system QoE and QoE fairness. These disclosed technology examples can estimate QoE without using application or customer data. Instead, they leverage temporal dependencies in network telemetry data to estimate QoE using machine learning models of application flows. Policies can govern the allocation of access categories and traffic priority to application flows based on the estimated QoE and can be dynamically adjusted to handle different classes of applications and varying network conditions.A policy agent can manage this allocation of access categories and traffic priorities. The policy agent can be trained in a simulation environment using the same knowledge base developed (collected) during the training of the QoE estimation models. This avoids the need to operate real clients, generate traffic through actual application sessions, and collect QoE and telemetry data.

[0009] In particular, examples of the disclosed technology are geared towards a resource manager that can be integrated into an existing network controller, such as a WLAN controller. A WLAN controller typically includes a management interface for configuring a Wi-Fi network, transmitting configurations to access points, and so on. The resource manager can determine an appropriate access category and traffic priority to assign to an application flow. The resource manager can use a probabilistic predictive machine learning model to estimate the QoE of an application flow based solely on telemetry data at a WLAN controller (WLAN controllers already receive telemetry data from network elements, such as access points, and are therefore a suitable network element for integration with the resource manager), although other network elements can also be used.Quality of Experience (QoE) metrics or characteristics are derived by application-specific machine learning models using a long short-term memory (LSTM) neural network. A reinforcement learning (RL) algorithm (using a double-deep Q-network (DDQN) RL method along with a feed-forward neural network) can then be used to determine the access category and traffic priority of each application flow based on the estimated QoE for that application flow. Parameters relating to the radio link / channel between client devices and an access point (AP) can be controlled by the WLAN controller. Application flows can be assigned to one of three access categories (best effort, video, or voice), and within each access category, the application flow can be assigned either a low or high priority (or other priority sets, if necessary) to control the packet leakage rate.The assignment of access category and traffic priority to an application flow to optimize QoE can be done in accordance with various possible combinations of access category and traffic priority, which are specified as an "action space".

[0010] It should be noted that the terms "optimize," "optimal," and the like, as used here, can be interpreted as making or achieving performance as effective or perfect as possible. However, as any professional reading this document will recognize, perfection cannot always be achieved. Accordingly, these terms can also mean making or achieving performance as good or effective as possible or practical under the given circumstances, or making or achieving performance better than that which can be achieved with other settings or parameters.

[0011] Before describing examples of the disclosed systems and methods in detail, it is useful to describe an exemplary network installation with which these systems and methods could be implemented in various applications. Fig. shows an example of a network configuration 100 that can be implemented for an organization, such as a company, an educational institution, a government agency, a health institution, or another organization. Fig. This shows an example of a configuration implemented in an organization with multiple users (or at least multiple client devices 110) and at least one physical or geographic location 102. The network configuration 100 can include location 102 in communication with a network 120.

[0012] Site 102 can include a primary network, which could be, for example, an office network, a home network, or another network installation. The primary network can be a private network, such as one that includes security and access controls to restrict access to authorized users. Authorized users might include, for example, employees of a company at site 102, residents of a house, or customers of a company.

[0013] In the example of Fig.Site 102 contains a controller 104 that communicates with network 120. The controller 104 ensures communication with network 120 for site 102. Besides the controller 104, there can be other communication points with network 120 for site 102. Although a single controller 104 is shown ( ), site 102 can include multiple controllers and / or multiple communication points with network 120. In some examples, the controller 104 can communicate with network 120 via a router. In other examples, the controller 104 provides router functionality to the devices at site 102. In this description, the word "tunnel" refers to an encapsulated mode for data transport between the access point and the controller.

[0014] Controller 104 can be operated to configure and manage network devices, for example, at site 102. Controller 104 can be operated to configure and / or manage switches, routers, access points, and / or client devices connected to a network. Controller 104 can itself be an access point (AP) or provide the functionality of one.

[0015] The Controller 104 can communicate with one or more Switches 108 and / or wireless APs 106A-C. Switches 108 and wireless APs 106A-C provide network connectivity for various client devices 110A-J. A client device 110A-J can access network resources, including other devices on the network (site 102) and on network 120, via a connection to a Switch 108 or AP 106A-C.

[0016] Examples of client devices include: desktop computers, laptops, servers, web servers, authentication servers, Authentication Authorisation Accounting (AAA) servers, Domain Name System (DNS) servers, Dynamic Host Configuration Protocol (DHCP) servers, Internet Protocol (IP) servers, Virtual Private Network (VPN) servers, network policy servers, mainframes, tablet computers, e-readers, netbook computers, televisions and similar displays (e.g., smart TVs), content receivers, set-top boxes, personal digital assistants (PDAs), mobile phones, smartphones, smart terminals, silent terminals, virtual terminals, video game consoles, virtual assistants, Internet of Things (IoT) devices, and the like.

[0017] Within site 102, switch 108 is included as an example of an access point to the network established at site 102 for client devices 110I-J. Client devices 110I-J can connect to switch 108 and access other devices within network configuration 100 via switch 108. Client devices 110I-J can also access network 120 via switch 108. Client devices 110I-J can communicate with switch 108 via a wired or wireless connection. In the example shown, switch 108 communicates with control unit 104 via a wired or wireless connection 112E.

[0018] The wireless APs 106A-C are another example of an access point to the network set up at site 102 for client devices 110A-H. Each of the APs 106A-C can be a combination of hardware, software, and / or firmware configured to provide wireless network connectivity for wireless client devices 110A-H. In the example of Fig. The APs 106A-C can be managed and configured by the controller 104. The APs 106A-C communicate with the control unit 104 and the network 120 via connections 112A-D, which can be either wired or wireless interfaces.

[0019] Network 120 can be a public or private network, such as the internet or another communications network, that enables connections between different locations, for example, location 102, and provides access to servers, such as server 130. Network 120 can include third-party telecommunications lines, such as telephone lines, broadcast coaxial cables, fiber optic cables, satellite communications, cellular communications, and the like. Network 120 can include any number of intermediate network devices, such as switches, routers, gateways, servers, and / or controllers, which are not directly part of Network Configuration 100 but facilitate communication between the various parts of Network Configuration 100 and between Network Configuration 100 and other entities connected to the network.

[0020] In one example, the resource manager mentioned above can be embodied in Server 130 and integrated with Controller 104 to optimally allocate network resources to client devices, thereby optimizing or maximizing overall QoE and QoE fairness. As mentioned above and described in more detail below, Server 130 (the resource manager) can: extract relevant data from telemetry data received from an access point; derive / estimate an application's unobservable QoE from this relevant telemetry data; and determine an optimal access category and traffic priority for application flows.

[0021] Fig.This shows a resource management system 200 and related operations for implementing adaptive network control according to some examples of the disclosed technology. In some examples, the resource management system 200 can contain one or more applications running on client devices or stations 210. An access point (AP), such as AP 206, can serve client devices 210. The resource manager 230 can, based on telemetry data from AP 206, estimate the quality of operation (QoE) for application flows passing through AP 206 based on previous or historical observations of the telemetry data. Furthermore, the resource manager 230 can learn an adaptive resource allocation policy that can be used to create or modify application flow configurations, which can be transmitted to AP 206 via a controller that manages operational aspects of AP 206, such as the WLAN controller 204. Finally, data can be transmitted to a data network, such as...the Internet 250, transmitted, received by it and routed through it.

[0022] As in Fig.As shown, data to and from client devices 210 can flow along the AP 206, to the WLAN controller 204, and to the internet 250 via a data path that includes connections 207 and 219. Telemetry data from AP 206 can be received by the telemetry processor 230A of the resource manager 230 via an SNMP (Simple Network Management Protocol) probe 204B of the WLAN controller 204. The SNMP probe 204B can use the SNMP protocol to query a device, such as AP 206, and obtain event data by acting as a trap daemon and monitoring SNMP traps and events, which in this example relate to raw telemetry data. Raw telemetry data can include user, system, and radio metrics or operational characteristics, such as... B. Radio link parameters between client devices 210 and AP 206, metrics relating to application traffic, wireless capabilities, the underlying Wi-Fi configuration and general system statistics.The raw telemetry data can be forwarded from the SNMP probe 204B to the telemetry processor 230A. The telemetry processor 230A can extract relevant data from the raw telemetry data and convert the data into the required formats if necessary. It should be noted that the relevant data extracted from the raw telemetry data is typically implementation-specific and may depend, for example, on the specific application or traffic type, network characteristics, and so on. Table 2, which is discussed in more detail below, contains example characteristics that are representative of such relevant data.

[0023] It is emphasized again that this telemetry data is the only data to which the Resource Manager 230 has access. Neither data from the Client Devices 210 nor the applications running on them are required or used by the Resource Manager 230.

[0024] A predictive machine learning (ML) model can be trained for each application class, and such predictive ML models can be loaded into the memory (not shown) of Resource Manager 230. It should be clear that the application class can refer to typical QoS classes for Layer 2 and Layer 3, such as a web browsing / email application class, a video conferencing application class, a YouTube® video streaming application class, an HD video streaming application class (Netflix®, Hulu®), and so on. The application classes are usually associated with the required throughput parameters, such as 500 Kbps to 1 Mbps for web browsing / email and 2-5 Mbps for HD video streaming, and so forth. In some examples, the application classes of interest are video streaming, video conferencing, and file transfer applications; however, any application class can be considered.A predictive machine learning (ML) model can be generated, trained, and loaded into the resource manager 230. These predictive ML models can be embodied by separate instances of the QoE estimator 230B. Thus, when data packets arrive at the WLAN controller 204 from the access point (AP) 206, the flow classifier 204A can identify a suitable application class for the data packets. The identified application class can be provided to the telemetry processor 230A. The telemetry data processed by the telemetry processor 230A can be forwarded to the corresponding application class instances of the QoE estimator 230B. The corresponding instance(s) of the QoE estimator 230B can estimate the QoE of the application flow reflected by the processed telemetry data, and the estimated QoE information or data can be forwarded to the policy agent 230C.

[0025] The Policy Agent 230C can calculate the access category and traffic priority for application flows based on estimated (determined as described above) QoE metrics to maximize overall QoE and QoE fairness. In some examples, RL can be used with multiple agents, where a specific Policy Agent or Policy Agent instance serves a particular application class, such as video streaming, video conferencing, or file transfer. The Policy Agent 230C instances can modify the application flow configurations associated with their respective application class via an AP Configuration Handler 204C, which then propagates these policies to the AP, such as AP 206.

[0026] The 204D data layer engine can handle the routing of data packets, for example via a table or other mechanism that can determine, for example, the destination path of an incoming data packet or the data path through the network.

[0027] It should be noted that the exchange of information about the access category and traffic priority between the Policy Agent 230C and the Configuration Handler 204C, the transmission of raw metric data (telemetry data) from the SNMP Probe 204B, and the transmission of application class information from the Flow Classifier 204A to the Telemetry Processor 230A can be carried out, for example, using the Transmission Control Protocol (TCP) and the necessary TCP interfaces between them. The transmission of the processed telemetry data from the Telemetry Processor 230A to the QoE Estimator 230B and the transmission of the application class-specific QoE from the QoE Estimator 230B to the Policy Agent 230C can be carried out via shared memory connections.As used here, a shared memory connection can refer to an area of ​​shared memory (not shown) that is used as a channel through which the telemetry processor 230A, the QoE estimator 230B, and the policy agent 230C can communicate with each other. Data exchange between the AP 206 and the WLAN controller 204 can take place via an OpenFlow interface.

[0028] Estimating Quality of Experience (QoE) presents challenges. Telemetry data is heterogeneous and can vary significantly in size and format. Furthermore, telemetry data may contain instantaneous values, cumulative statistics, and derived features, each exhibiting potential outliers and differing ranges, areas, and distributions. Additionally, the relationship between telemetry data and QoE is non-linear (and therefore complex) and relies on hidden or latent features. Ultimately, QoE metrics or features are often influenced by the temporal context and dependencies of the telemetry data. That is, an application's QoE may depend on the history of its telemetry data. Therefore, examples of the disclosed technology aim to accurately model the temporal and hidden relationship between telemetry and QoE metrics based on historical telemetry data, as mentioned above.

[0029] Fig.This shows an example of a computer component 300 that can be used to implement QOE estimation according to some examples of the disclosed technology. Regarding Fig. The computer component 300 could, for example, be a server computer, a controller, or another similar computer component capable of processing data. In the example implementation of Fig. The computer component 300 comprises a hardware processor 302 and machine-readable storage media 304. In some examples, the computer component 300 can be an embodiment of the resource manager 230 (or one or more components thereof, e.g., the QoE estimator 230B).

[0030] The hardware processor 302 can be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices capable of retrieving and executing instructions stored in machine-readable memory media 304. The hardware processor 302 can retrieve, decode, and execute instructions, such as instructions 306-316, to control processes or operations for application-flow-specific QoE estimation. Alternatively or in addition to retrieving and executing instructions, the hardware processor 302 can include one or more electronic circuits containing electronic components for performing the functionality of one or more instructions, such as a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other electronic circuits.

[0031] Machine-readable storage media, such as machine-readable storage media 304, can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. For example, machine-readable storage media 304 can be random access memory (RAM), non-volatile random access memory (NVRAM), electrically erasable programmable solid-state memory (EEPROM), a storage device, an optical disk, and the like. In some examples, the machine-readable storage media 304 can be non-transient, the term "non-transient" excluding transitive transmission signals. As detailed below, machine-readable storage media 304 can be encoded with executable instructions, such as instructions 306-316.

[0032] The hardware processor 302 can execute instruction 306 to vary the number of applications running simultaneously on client devices and enforce poor QoE.

[0033] In some examples of the disclosed technology, telemetry data can be processed sequentially to capture long-term dependencies while maintaining a reasonably long context over extended periods. Specifically, examples of the disclosed technology model QoE estimation using a Long Short Memory (LSTM) network, which can capture trends evolving over time and maintain representative gradient values ​​as the context length increases. It should be noted that QoE estimation can also be modeled using other machine learning (ML) model designs, but the predictive accuracy may vary.

[0034] To train an application-class-specific QoE estimation model according to the examples of the disclosed technology, applications can be run in a Wi-Fi test environment. Telemetry data and the precise QoE for video streaming, video conferencing, and file transfer can be collected according to the application execution. That is, an LSTM model can receive AP telemetry data as input, and the LSTM model can output the estimated QoE for the corresponding application class.

[0035] Telemetry data collection and quality of service (QoE) assessment can be performed under various network conditions, ranging from underutilized to overloaded scenarios, by systematically varying the Wi-Fi configurations. Table 1 provides an example of possible Wi-Fi control parameters that can be used to generate different Wi-Fi configurations. Furthermore, the number of concurrently running applications can be varied from 1 to n, where n is chosen such that all concurrently running application instances overload the network, resulting in poor QoE. For example, application sessions can be evenly distributed across multiple Wi-Fi client devices, such as four client devices. Table 1 Frequency band PHY / MAC mode Channel width Category Access Traffic priority 2.4 GHz High throughput (HT) 20 MHz Background(BK) High Priority (HQ) 5 GHz Very HT (VHT) 40 MHz Best Effort (BE) Low Priority (LQ) High efficiency (HE) 80 MHz Video (VI) 160 MHz Vote (VO)

[0036] The 302 hardware processor can execute the 308 command to run applications on client devices. Applications can be run for each application session; in this example, four client devices. The duration of an application session can vary, but in one example, an application session can be set to 200 seconds. Application execution can be repeated many times, for example, twice in some examples, to increase the size and variety of the dataset.

[0037] The 302 hardware processor can execute instruction 310 to collect telemetry data associated with application execution and measure the QoE of applications on client devices. Using the application execution approach described above, extensive telemetry data can be collected. For example, applications running for a total of 251 hours across all client devices for the three application classes (video streaming, video conferencing, and file transfer applications) can generate over 226,000 data points. The telemetry data can be collected according to a time-based specification, such as every four seconds.

[0038] Furthermore, QoE can be measured during application execution on client devices. Application QoE metrics can be calculated at any interval and normalized against a maximum achievable QoE value for each application. This allows for comparison of QoE values ​​across application classes. While subjective metrics are typically the most accurate indicators of a user's experience, capturing them is impractical across a large number of system configurations. Therefore, objective quality metrics can be used to adequately reflect the quality perceived by an end user.

[0039] To calculate the Quality of Experience (QoE) of a video stream, algorithms based on Model Predictive Control (MPC) can be used, where the QoE metric is a function of the following factors: the quality of individual video chunks, quality variations between successive chunks, replay time, and start delay. The underlying principle of this approach is that users prefer uninterrupted video playback (without interruption) once a video starts playing, and that frequent changes in video quality between segments are undesirable. MPC achieves the best possible QoE during a video streaming session.

[0040] To measure the quality of video conferences, the focus is on the following metrics: audio jitter, audio packet loss, video packet loss, and transmitted video packets. The reason for considering these metrics is that end users generally find choppy real-time sound or video images that are out of sync with the audio to be disruptive and undesirable. The best possible Quality of Experience (QoE) for the video conferencing application is achieved with the highest video resolution, the maximum possible frame rate, and minimal jitter and packet loss. For video conferencing applications, the QoE values ​​are normalized to this best possible QoE value. To map these QoE metrics to a four-second interval, the QoE values ​​can be averaged over four seconds, and a single QoE value can be reported.

[0041] Since the QoE for file transfers is calculated at each time interval (as described above) and not solely based on the file transfer completion time, the number of bytes transferred from the server to the client per second can be recorded. This data can then be averaged over a four-second interval to calculate the QoE for that period, comparing it to the established four-second interval for telemetry collection. The rationale for calculating the QoE of file transfers in this way is that end users prefer file transfers to complete in the shortest possible time, which can be achieved if a server consistently / continuously sends large amounts of data to a recipient.

[0042] The 302 hardware processor can execute instruction 314 to reconcile application-running telemetry data with measured application-flow-specific QoE. In some examples, the QoE and telemetry data are synchronized based on their timestamps. Since each application reports its QoE at a different rate than the telemetry data, these measurements can be reconciled by averaging the QoE metrics over each four-second interval. This ensures that the telemetry data accurately reflects the system state at time t by mapping the average QoE metric from t to t seconds onto the corresponding telemetry.

[0043] Furthermore, telemetry data harmonization can also include data normalization. In the case of video streaming applications, the QoE of each (mentioned above) video chunk can be normalized to the best possible QoE value. QoE values ​​range from negative infinity to one, although experimentally, QoE generally does not go below minus one. Therefore, the range of QoE values ​​can be restricted to between minus one and one. These QoE values ​​between zero and one can be normalized using a min-max normalization method. These normalized QoE values ​​are reported for each video chunk, corresponding to a four-second interval. It is important to note that a video streaming application runs in balanced mode to achieve a balance between avoiding re-buffering (stalls) and avoiding instability (changes in video quality).This results in minimal buffering and stable quality.

[0044] In video conferencing applications, the best possible QoE is generally achieved with the highest video resolution, the maximum possible frame rate, and minimal jitter and packet loss. In video conferencing applications, the QoE values ​​are normalized to this best possible QoE value.

[0045] For file transfer applications, the best achievable QoE for file transfer can be determined with an optimal Wi-Fi configuration that includes a 160 MHz channel in the 5 GHz frequency band and a physical HE interface. With this configuration, the aforementioned server can transfer the maximum number of bytes to the client device, representing the best achievable QoE. The QoE for each four-second interval can then be normalized based on this maximum QoE value for file transfer.

[0046] The 302 hardware processor can execute instruction 314 to train a QoE model, keeping in mind that a QoE machine learning (ML) model can be developed and implemented for each application class. To train an application-class-specific QoE model, an LSTM mesh can be constructed with an input layer consisting of, for example, 257 telemetry data points and three stacked LSTM layers, where the first, second, and third layers have 256, 128, and 64 hidden units, respectively. A dense output layer can follow the LSTM layers to compute the estimated QoE metric. The application-class-specific QoE model can be trained using the Adam optimizer and the mean squared error (MSE) loss function. The dataset can be split into training, validation, and test datasets, using split ratios of 60%, 20%, and 20%, respectively.The training process can span 50 epochs with a stack size of 32. Table 2 (below) shows examples of features identified by the model across all applications using the permutation feature meaning method for QoE estimation. Table 2 feature Description delay End-to-end delay RX Unicast Data Frames Number of received unicast frames MCS-X transmit data frame Total number of data frames transmitted at the MCS rate X Last SNR Last recorded signal-to-noise ratio Last RX SNR SNR of the last data packet received by the client Last ACK SNR SNR of the last acknowledgment packet sent by the client Current noise level Remaining background noise that is detected by an access point Channel occupied 1 / 4 / 64 sec. Percentage of time the radio channel was occupied in the last 1 / 4 / 64 seconds PS condition Power saving state, indicating whether the AP / channel / link is awake or in power saving state. TX retry attempts Number of packets that the AP had to resend to the client due to a transmission error. TX RTS failed Number of RTS (Ready To Send) frames that were not successfully transmitted Health Quality of the connection between the customer and the radio device

[0047] Although the state of the PO-MDP is estimated using the QoE estimator 230B, the transition probability remains difficult to model due to the complex relationship between the actions performed and the observed state. Examples of the disclosed technology address the complexity of using RL for a large search space and the potential lack of sufficient training data by employing an RL approach that can explore various possibilities within the search space while efficiently utilizing the collected state. The exploration process can be controlled to accelerate model convergence, while a DNN handles the high-dimensional state-action space in the Wi-Fi environment, which is caused by different application classes, dynamic network conditions, and varying client capabilities.As mentioned previously, the Policy Agent 230C can utilize a DDQN-RL method in conjunction with a neural feed-forward network to handle the high-dimensional space, as well as a repetition buffer to improve sampling efficiency.

[0048] In particular, the Q-network, denoted Q(s, a;g θ), is a neural network that approximates the action value function, where s represents the state, a the action, and θ the network parameters. The target Q-network Q' is a copy of the Q-network that is regularly updated to stabilize the learning process and minimize any potential overestimation of the Q-values. During training, the Q-value is updated as follows: Q(st,at)←Q(st,at)+α[rt+γQ'(st+1,argmaxa,Q(st+1,a';θ)θ')−Q(st,at;θ)],

[0049] Whereas s t the current state at time t is, a tthe action performed at time t, a' the action that maximizes the Q-value of the next state, r t the reward received at that time is, s t where +1 is the next state at time t+1, θ' is the target parameter of the Q-network, a is the learning rate, and y is the discount factor. Actions can be selected using Boltzmann exploration, which assigns probabilities to each action based on its Q-values. P(a)=eQt(st,a) / T∑i=1neQt(st,a') / T, where T is the temperature parameter that controls the trade-off between exploration and exploitation.

[0050] The loss function used to train the Q-network is the mean squared error (MSE) between the predicted Q-values ​​and the target Q-values: L(θ)=E[(rt+γQ'(st+1,argmaxa'Q(st+1,a';θ)θ')−Q(st,at;θ))2].

[0051] As previously mentioned, examples of the disclosed technology control the parameters of the radio link between Wi-Fi client devices and the access point (AP). Typical WLAN controllers allow the configuration of parameters at the flow, radio, and device levels. Since configuring parameters at the radio and device levels requires a restart of the AP or the devices, Resource Manager 200 configures flow-level parameters that can be modified at runtime. This runtime modification, in accordance with examples of the disclosed technology, provides the ability to achieve a desired improvement in the quality of service (QoE) of a flow while maintaining QoE fairness across, for example, multiple flows. Such parameters can be calculated by Resource Manager 200 and ultimately modified via the WLAN Controller 204's configuration handler 204C.

[0052] Access category and traffic priority can be assigned to an application flow individually or in combination. When assigning the access category aspect to an application flow, the application class of the application flow does not need to be restricted to a specific access category. Instead, the Policy Agent 230C can assign incoming application streams to any access category based on the learned policy to adapt to application heterogeneity and varying network conditions. As mentioned above, in some examples, application streams can be assigned to one of three access categories: Best Effort (BE), Video (VI), and Voice (VO). Furthermore, the Policy Agent 230C can assign either a low priority (LQ) or a high priority (HQ) to each stream, thereby controlling the packet outflow rate. These actions (i.e.,The assignment of the AC / traffic priority combination for each application can be performed at a specific frequency, e.g., every 16 seconds. Such periodic execution of actions allows the Policy Agent 230C (and the Resource Manager 230 as a whole) to dynamically adapt to changing network conditions. The duration / frequency of execution can vary, e.g., depending on the processing capacity and speed of the WLAN Controller 204's processor (in ). Fig. (not shown).

[0053] Fig.This shows an example action space 400 that defines combinations of access categories / traffic priorities to which incoming application flows can be assigned, where an "action," as used here, can refer to the process of assigning possible combinations of access categories / traffic priorities to application flows. This is achieved by combining low- and high-priority (LQ / HQ) queues with the access categories, as shown in Fig.As illustrated, according to an example of the disclosed technology, action space 400 comprises the following six possible action combinations: LQ+BE and HQ+BE for BE access category 402; LQ+VI and HQ+VI for VI access category 404; and LQ+VO and HQ+VO for VO access category 406. Incoming data packets representative of one or more application streams 408 can be assigned by the Policy Agent (e.g., Policy Agent 230C) to a specific combination of access category / traffic priority defined in action space 400. By assigning an access category / traffic priority to an application flow at regular intervals (e.g., every 16 seconds), as described above, the Policy Agent modifies / updates the application flow configurations via Configuration Handler 204C. Configuration Handler 204C can then send these policies to the AP, e.g., B. the AP 206, through which the application streams flow.

[0054] It should be noted that the smallest unit at which a policy agent takes action (referred to as a state space) is a 5-tuple application flow traversing an AP. Given the complexity and scope of the data available in the controller database, which includes thousands of metrics, determining which of these metrics significantly impact QoE estimation can be challenging. To ensure comprehensive coverage, all available metrics that allow the QoE estimator to select a relevant subset and accurately estimate QoE are considered, in accordance with the disclosed technology examples.

[0055] AP telemetry data can be categorized into three data types: application-specific data, client-specific data, and access point-specific data. Application-specific data may include metrics such as flow-specific throughput, flow transport protocol, and TX / RX bytes. Client-specific data may include the client's wireless capabilities, such as signal-to-noise ratio (SNR), PHY / MAC interface details, channel width, and access category. Access point-specific data may include system-specific metrics, such as total packet loss, lost packets, buffer SX / RX metrics, and various radio metrics. The metrics listed above are merely examples to illustrate the context, as many other relevant metrics may be considered within each data category.

[0056] As mentioned above, a policy agent configured according to the disclosed technology examples uses a DDQN RL procedure coupled with a neural feed-forward network. In some examples, the DDQN's reward function is configured to maximize overall QoE and QoE fairness. Due to the aforementioned use of a multi-agent RL (remember that different policy agent instances serve different application classes), each policy agent instance receives, in addition to state: (1) the average QoE of all other applications; and (2) the deviation of that application's QoE from the average QoE. After an action in state t Once the process has been carried out, the reward function can be defined as follows: R(stapp,atapp)=wQoE∗QoEapp−wmean∗QoE¯others−wdeviation∗deviation,

[0057] Where QoE appThe QoE of a current application flow is QoE others The average QoE of all other application flows is and deviation = |QoE app - QoE system | the deviation of the QoE of the current application from the system-wide, average QoE, where QoE¯system=1N∑i=1NQoEi The system-wide average QoE for N application flows is [missing information]. The weights w QoE , w mean and w deviation These settings are used to adjust the sensitivity of the application's QoE to the Wi-Fi environment, allowing the reward function to be tailored to the specific needs of different applications. This function maximizes the QoE of the current application, compares it to the average QoE of all other applications, and maximizes QoE fairness by minimizing deviations in QoE metrics.

[0058] Regarding the resource manager's policy agent, training it in Maestro presents several challenges. The training process involves accessing state variables, executing actions, and observing rewards and subsequent states. Achieving this in a real-world, working Wi-Fi facility or environment requires running actual clients, generating traffic through real application sessions, and collecting both QoE and telemetry data. This means that, due to the complexity and variability of real-world network conditions, it's not feasible to obtain a sufficient number of diverse samples from which the policy agent can learn an allocation policy for access categories / traffic priorities.Furthermore, these real-world Wi-Fi conditions cannot be easily integrated into current RL environments, not to mention that exploring all possible state-action combinations is not feasible, as it would require an excessive expenditure of time and resources.

[0059] Therefore, according to an example of the disclosed technology, a simulation environment that utilizes an extensive knowledge base already gathered during the training of the QoE estimator can be used to train the policy agent. This knowledge base includes telemetry data and the corresponding QoE values, enabling the resource manager to efficiently simulate the training process for the policy agent.

[0060] Fig.Figure 500 shows an example of Algorithm 500, which represents the iterative learning process used by the policy agent in a simulation environment. The strategy agent starts with a random state in the knowledge base. After selecting an action based on this state, Algorithm 500 evaluates the action by comparing it to the knowledge base. Algorithm 500 uses the Pearson correlation coefficient ρ to find the best-matching vector v. By maximizing ρ, Algorithm 500 identifies the best-matching vector v that corresponds to the current state, the Wi-Fi configuration, the action taken, and the reward received. This matching vector, v, represents the next state, allowing the strategy agent to continue its iterative learning process.

[0061] To estimate the weights of the reward function in Eq. 4 and the temperature value of the Boltzmann exploration in Eq. 2, five different representative sets of reward function weights and ten temperature values ​​in the range from 0 to 1 with increments of 0.1 can be used. A combination of these values ​​can each be fed individually into Algorithm 500. The final weight and temperature values ​​can then be selected based on the loss function.

[0062] The loss function of RL training behaves similarly to the function (Lt) = α * t * e - βt, where t is the training time or the number of episodes, α is a scaling factor that determines the height of the peak, and β is a factor that controls the rate of decline after the peak. This is because the function represents the expected loss values ​​over time. The loss value initially increases as the strategy agent explores, reaches a peak, and then gradually decreases to zero as the strategy agent learns and optimizes its policies (access category / traffic priority assignment). Reward weights and temperature values ​​can be chosen by fitting this function to the observed loss values.

[0063] Table 4 contains an example set of optimized reward weights and temperature values ​​for each application class, based on a state size of 257 and an action size (space) of six, a hidden layer size of 128, a discount factor of 0.99, a stack size of 32, an Adam optimizer, a learning rate of 0.0001, and a repetition buffer of 10,000. Table 4 Registration w QoE w mean w deviation T Video streaming 0.2 0.6 0.2 0.4 Video conferences 0.2 0.2 0.6 0.6 File transfer 0.34 0.33 0.33 0.8

[0064] Fig. This shows an example of a computer component 600 that can be used to implement QOE-capable dynamic resource allocation in accordance with some examples of the disclosed technology. Referring to Fig. Computer component 600 (similar to computer component 300) can, for example, be a server computer, a controller, or another similar computer component capable of processing data. In the example implementation of Fig.Computer component 600, like computer component 300, comprises a similar hardware processor 602 and a similar machine-readable storage medium 604. In some examples, computer component 600 may be an embodiment of resource manager 230, a server such as server 160, etc.

[0065] As described in detail below, machine-readable storage media 604 can be encoded with executable instructions, e.g. with instructions 606-612.

[0066] The 602 hardware processor can execute the 606 instruction to receive telemetry data from an access point (AP) at a network controller. As previously discussed, client devices can be served by an AP while running applications such as video streaming, file transfer, and the like. Unlike other systems and methods that require collecting, for example, internal application state information or other data from the client device(s) or application(s), examples of the disclosed technology are able to estimate Quality of Energy (QoE) solely based on telemetry data associated with the execution of applications on the client devices whose data packets / traffic traverse the AP. The telemetry data can be parsed, as described above, to extract relevant data, and it can be normalized / transformed to account for variations between data types.

[0067] The 602 hardware processor can execute instruction 608 to estimate a QoE value for individual application traffic streams of the one or more application traffic streams passing through the AP, based on telemetry data. A telemetry processor, which receives telemetry data from the AP and performs the aforementioned extraction of relevant data and data normalization, can forward the relevant data to a QoE estimator. The QoE estimator can comprise multiple QoE estimator instances used to estimate a QoE for the application(s) running on the client device(s), as evidenced by the received telemetry data according to an application class into which the application(s) are categorized. The QoE estimator instances comprise machine learning models trained to accurately model the temporal and hidden relationship between telemetry data and QoE metrics.This means that the QoE estimator instances derive "unobservable" QoE metrics. In some examples, an LSTM network can be constructed to train these QoE ML models, where the training data is derived from telemetry data and measured QoE information under different network conditions, and where the telemetry data and the measured QoE information are compared, for example, using timestamps.

[0068] The 602 hardware processor can calculate an access category and traffic priority based on the estimated QoE value for each application traffic stream and assign them to that stream. In some examples, multiple policy agents are trained with machine learning models, which in turn are targeted to specific application classes, to maximize overall QoE and QoE fairness within the network and among AP / client devices. Reinforcement learning can be used for this assignment. Ultimately, the application-class-targeted policy agent assigns an action to the application streams belonging to its application class. This action includes assigning both the access category and traffic priority to the application streams. It should be noted that application streams can be assigned to any access category based on the policies learned by the policy agent instances.Furthermore, it should be noted that the policy agent instances can be trained in a simulated environment to avoid the operational costs associated with running actual client devices and applications, as well as the complex operation of a Wi-Fi network. The simulation environment can utilize existing knowledge bases containing received telemetry data and corresponding QoE values.

[0069] Hardware processor 602 can execute instruction 612 to maximize the overall QoE across one or more application traffic streams and the QoE parity applied to each individual traffic stream by adjusting the configurations of each traffic stream according to its assigned access category and traffic priority. As mentioned above, in some examples, configuration actions can be performed at a specific frequency to accommodate dynamic network conditions. A configuration handler ensures that the policies / configurations discovered / assigned by the policy agent instances are transmitted to the access point.

[0070] Fig.Figure 700 shows a block diagram of an example Computer System 700, in which various examples of the technology described here can be implemented. The Computer System 700 comprises a bus 702 or other communication mechanism for transmitting information, and one or more hardware processors 704 connected to the bus 702 for processing information. The hardware processor(s) 704 could, for example, be one or more general-purpose microprocessors.

[0071] The Computer System 700 also includes a main memory 706, such as random access memory (RAM), a cache, and / or other dynamic memory devices connected to the 702 bus to store information and instructions to be executed by the 704 processor. The main memory 706 can also be used to store temporary variables or other intermediate information during the execution of instructions to be carried out by the 704 processor. Such instructions, stored in memory media accessible to the 704 processor, make the Computer System 700 a specialized machine, adapted to perform the operations specified in the instructions.

[0072] The Computer System 700 also includes a read-only memory (ROM) 708 or other static storage device connected to bus 702 to store static information and instructions for the processor 704. A storage device 710, such as a magnetic disk, an optical disk, or a USB flash drive, etc., is provided and connected to bus 702 to store information and instructions.

[0073] The Computer System 700 can be connected via bus 702 to a display 712, such as a liquid crystal display (LCD) (or a touchscreen), to show information to a computer user. An input device 714, including alphanumeric and other keys, is coupled to bus 702 to transmit information and command selections to the processor 704. Another type of user input device is the cursor control 716, such as a mouse, trackball, or cursor direction keys, for transmitting directional information and command selections to the processor 704 and for controlling cursor movement on the display 712. In some examples, the same directional information and command selections as with cursor control can be implemented by receiving touch inputs on a touchscreen without a cursor.

[0074] The Computer System 700 can include a user interface module for implementing a graphical user interface, which can be stored on a mass storage device as executable software code that is executed by the computer device(s). This and other modules can include components such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.

[0075] Elements that comprise a QoE-conscious dynamic resource allocation system, e.g., that of Fig. , including, for example, a resource manager, a network orchestrator, etc., can be embodied by the Computer System 700

[0076] The term "non-volatile media" and similar terms as used here refer to all media that store data and / or instructions that cause a machine to operate in a particular way. Such non-volatile media can include both non-volatile and volatile media. Examples of non-volatile media include optical or magnetic disks, such as the Storage Device 710. Examples of volatile media include dynamic storage devices, such as the Main Memory 706. Non-volatile media are distinct from transmission media but can be used in conjunction with them. Transmission media are involved in the transfer of information between non-transmittable media.

[0077] The Computer System 700 also includes a 718 communication interface, which is connected to the 702 bus. The 718 network interface provides a two-way data communication connection to one or more network connections, which are connected to one or more local area networks. The 718 communication interface can be, for example, an ISDN (Integrated Services Digital Network) card, a cable modem, a satellite modem, or a modem to establish a data communication connection to a corresponding type of telephone line. Wireless connections are also possible. In each of these implementations, the 718 network interface sends and receives electrical, electromagnetic, or optical signals that carry digital data streams with various types of information.

[0078] Each of the processes, methods, and algorithms described in the preceding sections can be embodied in code components and fully or partially automated by them, which are executed by one or more computer systems or computer processors with computer hardware. The one or more computer systems or computer processors can also be operated in such a way as to support the execution of the corresponding operations in a cloud computing environment or as Software as a Service (SaaS). The processes and algorithms can be partially or fully implemented in application-specific circuits. The various features and procedures described above can be used independently or combined in various ways.Various combinations and subcombinations fall within the scope of this disclosure, and certain procedural or process blocks may be omitted in some implementations. The methods and processes described herein are also not restricted to a particular order, and the associated blocks or states may be executed in other suitable orders, in parallel, or otherwise. Blocks or states may be added to or removed from the disclosed examples. The execution of certain operations or processes may be distributed across computer systems or computer processors that are not located in a single machine but are distributed across a number of machines.

[0079] The term "or" used here can be understood in both an inclusive and an exclusive sense. Furthermore, the description of resources, processes, or structures in the singular should not be interpreted as excluding the plural. Conditional expressions such as "can," "could," "might," or "may" generally indicate that certain examples include specific features, elements, and / or steps, while other examples do not, unless explicitly stated otherwise or understood differently in the given context.

[0080] Unless explicitly stated otherwise, the terms and expressions used in this document, as well as their variations, are to be understood as open rather than restrictive. Adjectives such as "conventional," "traditional," "normal," "standard," "known," and terms with similar meanings are not to be understood as limiting the described subject matter to a specific period or to an item available at a particular time, but should be understood as encompassing conventional, traditional, normal, or standard technologies that may be available or known now or at any time in the future.The presence of expansive words and phrases such as "one or more", "at least", "but not limited to" or similar phrases in some cases is not to be understood as implying that the narrower case is intended or required when such expansive phrases are not present.

Claims

[1] Procedure comprising the following: to receive telemetry data from an access point (AP) from a network controller of a network; to estimate a Quality-of-Experience (QoE) value for individual application traffic streams of one or more application traffic streams passing through the AP based on telemetry data; to calculate and assign an access category and traffic priority for each application traffic flow based on the estimated QoE value; Maximizing the overall QoE across one or more application traffic streams and the QoE parity as applied to each application traffic stream by adjusting the configurations of each application traffic stream in accordance with their respective assigned access category and traffic priorities. [2] Method according to claim 1, wherein the received telemetry data comprises raw telemetry data including user metrics, network metrics of the network and radio metrics from one or more radio devices operated in the AP. [3] The method according to claim 2 further comprises processing the raw telemetry data to extract telemetry data relevant for QoE estimation. [4] Method according to claim 3, wherein the processing of the raw telemetry data comprises the processing of sequential telemetry data. [5] Method according to claim 1, wherein the estimation of the QoE value comprises identifying an application class that is associated with packets of the one or more application traffic streams received by the network controller. [6] Method according to claim 5, further comprising the use of the identified application class to apply an application class-specific prediction model corresponding to the identified application class to estimate the QoE value based on the telemetry data. [7] Method according to claim 6, wherein the application class-specific prediction model comprises a long short-term memory (LSTM) neural network. [8] Method according to claim 6, further comprising training the application class-specific prediction model using collected telemetry data and measured QoE under different network conditions, including underloaded and overloaded network conditions. [9] Method according to claim 8, further comprising synchronizing the collected telemetry data with the measured QoE on the basis of respective timestamps associated with the collected telemetry data and the measured QoE. [10] Method according to claim 1, wherein the calculation and assignment of an access category and a traffic priority to the individual application traffic flows is performed on the basis of the estimated QoE value of application class-specific policy agents. [11] Method according to claim 1, wherein the application class-specific policy agents use a Double Deep Q-Network (DDQN) Reinforcement Learning (RL) algorithm coupled with a neural feed-forward network. [12] Method according to claim 11, wherein the calculation and assignment of an access category and a traffic priority to the individual application traffic flows based on the estimated QoE value comprises the joint assignment of the access category and the traffic priority in accordance with an action space that includes possible combinations of access category and traffic priority. [13] Method according to claim 10, further comprising training the application class-specific strategy agents in a simulation environment. [14] System comprising the following: a processor; and a memory unit that contains instructions which, when executed, cause the processor to: Receive telemetry data from an access point (AP); to estimate an Experience Quality of Experience (QoE) score for each application traffic stream of a multitude of application traffic streams passing through the AP based on telemetry data; Calculate an access category and traffic priority for each application traffic stream based on the estimated QoE value; the calculated access category and traffic priority common to each application flow; and to push a configuration that includes the jointly assigned access category and traffic priority to a network controller so that it can be forwarded to the AP to reconfigure the AP to process subsequent data packets belonging to each of the application traffic streams in accordance with the jointly assigned access category and traffic priority. [15] System according to claim 14, wherein the instructions which, when executed, cause the processor to estimate the QoE value include instructions which, when executed, also cause the processor to: Identification of an application class associated with data packets of one or more application traffic streams received by the network controller; and Application of an application class-specific prediction model that corresponds to the identified application class to estimate the QoE value based on the telemetry data. [16] System according to claim 15, wherein the application class-specific prediction model comprises a long short-term memory (LSTM) neural network. [17] System according to claim 14, wherein the memory unit contains further instructions which, when executed, cause the processor to train the application class-specific prediction model using collected telemetry data and measured QoE under various network conditions, including underloaded and overloaded network conditions. [18] System according to claim 14, wherein the instructions which, when executed, cause the processor to calculate and jointly assign the access category and traffic priority, comprise further instructions which, when executed, cause the processor to execute specific policy agent instances for each identified application class. [19] System according to claim 18, wherein the agents of the policy agent instances use a Double Deep Q-Network (DDQN) Reinforcement Learning (RL) algorithm coupled with a feed-forward neural network to compute the access category and traffic priority. [20] System comprising the following: a processor; and a memory unit that contains instructions which, when executed, cause the processor to: Running a number of applications simultaneously on client devices to force a poor Quality of Experience (QoE) for client devices served by an Access Point (AP) in a network; to run the applications on the client devices; Collect telemetry data from the AP and measure QoE, with the telemetry data being linked to the execution of the applications; and Training a QoE model with the telemetry data and the measured QoE, which is to be operationalized to maximize the overall QoE across a multitude of traffic flows passing through the AP, and the QoE parity applied to individual members of the multitude of traffic flows, where each member of the multitude of traffic flows is associated with one of the number of applications running on the client devices.