Scheduling radio resources in a communications network
By acquiring users' physiological parameters, using machine learning models to predict latency perception thresholds, and combining this with reinforcement learning to schedule radio resources, the problem of existing technologies failing to consider human user behavior and mental state is solved, achieving efficient resource utilization and improved user experience quality.
Patent Information
- Application Number
- CN202080101427.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-28
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-05-28
AI Technical Summary
Existing radio resource allocation algorithms fail to effectively consider the behavior and mental state of human users, leading to resource waste and network inefficiency, especially in highly interactive applications such as virtual reality and immersive games.
By acquiring users' physiological parameters, such as heart rate and blood pressure, machine learning models are used to predict users' latency perception thresholds. Reinforcement learning is then combined to schedule radio resources in order to provide services with a user-perceived, pre-defined quality of experience.
It improves the utilization efficiency of radio resources, reduces network resource waste, achieves user-perceived QoS and power savings, adapts to users' cognitive and mental states, and enhances network energy efficiency.
Smart Images

Figure CN115699964B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to methods, nodes, and systems in communication networks. More specifically, but not exclusively, this disclosure relates to scheduling resources in communication networks. Background Technology
[0002] User-centric applications such as virtual reality and immersive gaming will see increasingly widespread adoption in future wireless networks. Common characteristics of these services include: a) a high level of interaction between the user and the application; and b) a greater demand for network resources compared to traditional cellular applications. Providing resources for such intensive applications presents a challenge to network resources. The purpose of the embodiments described herein is to improve resource provisioning in communication networks, particularly for resource-intensive applications involving high levels of user interaction. Summary of the Invention
[0003] As mentioned above, emerging user-centric applications such as virtual reality and immersive gaming will place increasing pressure on communication networks. Existing radio resource allocation algorithms allocate radio resources based on channel and network conditions. For example, a base station (BS) may seek to allocate resource blocks (RBs) and power to users based on the BS's latency requirements and channel conditions, ignoring user behavior and state. Traditional algorithms still rely on device-level characteristics and cannot be aware of human end-users and their characteristics (e.g., cognitive limitations or behavior). Therefore, traditional algorithms may waste network resources by allocating more resources to human users who, for example, cannot perceive the associated QoS gains due to cognitive limitations. Therefore, when deploying user-centric applications on wireless and cellular systems, resource scheduling can be improved by making the network aware not only of the application's Quality of Service (QoS) requirements but also of the human user's perception of that QoS (see the paper "Human-in-the-Loop Wireless Communications: Machine Learning and Brain-Aware Resource Management" by A. Kasgari, W. Saad, and M. Debbah).
[0004] In their paper, “Prospect Pricing in Cognitive Radio Networks,” Y. YANG, L. Park, N. Mandayam, I. Seskar, A. Glass, and N. Sinha, the authors conducted experiments comparing subjective and objective measurements of video quality. For each selected pair of grouped loss and delay, they objectively measured (using video frames decoded per second) the corresponding frames decoded per second at the video player used to display the video. Psychophysical experiments have revealed that, among the selected parameters, video frames decoded per second is the best objective indicator of video quality, while the perception of the number of interruptions and stutters is the best subjective indicator of overall video quality. Human subjects were also asked to subjectively rate the quality of the video on a four-level scale, where 4 is the highest rating and 1 is the lowest. The results show that the relationship between objective and subjective probabilities exhibits an inverse S-shaped probability weighting effect.
[0005] In the embodiments of this paper, an AI-assisted brain perception resource management process is proposed, wherein the proposed resource allocation method takes into account human behavior and mental state while considering channel state information.
[0006] In a first aspect, a computer-implemented method is provided for scheduling radio resources to a user equipment (UE) in a node of a communication network to provide services to a user of the UE. The method includes: acquiring one or more physiological parameters of a user of the UE; and scheduling resources to the UE based on the one or more physiological parameters to provide services to the UE with a predetermined quality of experience perceived by the user.
[0007] Physical parameters can be correlated with a user's alertness. Therefore, by considering a user's physiological parameters when allocating resources in a communication network, resources can be allocated in a way that takes human behavior and mental state into account (e.g., while considering network parameters such as channel state information). This information can be obtained and used transparently, for example, by reminding the user of their physiological parameters or requiring their permission to use it. Thus, users can be aware of the effects of using the disclosed methods and how the methods actually work.
[0008] According to a second aspect, a node in a communication network is provided for scheduling radio resources to a user equipment (UE) to provide service to a user of the UE. The node includes: a memory including instruction data representing an instruction set; and a processor configured to communicate with the memory and execute the instruction set. When executed by the processor, the instruction set causes the processor to acquire one or more physiological parameters of the user of the UE, and schedule resources to the UE based on the one or more physiological parameters to provide service to the UE with a predetermined quality of experience perceived by the user.
[0009] According to a third aspect, a computer program product including a computer-readable medium is provided, the computer-readable medium having computer-readable code contained therein, the computer-readable code being configured to cause the computer or processor to perform the method of the first aspect when executed by a suitable computer or processor. Attached Figure Description
[0010] To better understand and more clearly illustrate how the embodiments described herein can be implemented, reference will now be made to the accompanying drawings by way of example only, in which:
[0011] Figure 1 Nodes in a communication network according to some embodiments are shown;
[0012] Figure 2 Methods in nodes of a communication network according to some embodiments are illustrated; and
[0013] Figure 3 A method in a node of a communication network according to some embodiments is shown. Detailed Implementation
[0014] In some situations, the human brain may not be able to perceive any differences between videos transmitted at different QoS levels (e.g., rate or latency). In order to deliver immersive, human-centered services, networks must adapt the use and optimization of wireless resources to the inherent characteristics of their human users, such as their behavior and brain processing limitations, thereby making more efficient use of available radio resources.
[0015] As described above, the embodiments of this document relate to scheduling resources in a communication network based on the physiological parameters of the end user. For example, physiological parameters can be used as indicators of the brain's attention and perception. In this way, for example, more resources can be allocated if the user is alert and has high cognitive processing compared to a case where the user has low cognitive processing and therefore does not perceive the increased QoS associated with any additional resources.
[0016] The embodiments described herein relate to communication networks. Generally, a communication network (or telecommunications network) may include any one or any combination of the following: wired links (e.g., ASDL) or wireless links (e.g., Global System for Mobile Communications (GSM), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), WiFi, or Bluetooth wireless technologies). Those skilled in the art will understand that these are merely examples and communication networks may include other types of links. Wireless networks can be configured to operate according to specific standards or other types of predefined rules or procedures. Therefore, specific embodiments of wireless networks may implement: communication standards such as Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Long Term Evolution (LTE), and / or other suitable 2G, 3G, 4G, or 5G standards; wireless local area network (WLAN) standards such as the IEEE 802.11 standard; and / or any other suitable wireless communication standards such as Global Microwave Access Interoperability (WiMax), Bluetooth, Z-Wave, and / or ZigBee standards.
[0017] Figure 1 Nodes in a communication network according to some embodiments of this document are illustrated. Typically, node 100 may include any component or network function (e.g., any hardware or software module) in the communication network suitable for performing the functions described herein.
[0018] For example, in some embodiments, a node may include a device capable of, configured to, arranged to, and / or operable to communicate directly or indirectly with a UE (e.g., a wireless device) and / or with other network nodes or devices in the communication network to enable and / or provide wireless or wired access to the UE and / or perform other functions (e.g., management) in the communication network. Examples of nodes include, but are not limited to, access points (APs) (e.g., radio access points), base stations (BSs) (e.g., radio base stations, NodeBs, evolved NodeBs (eNBs or eNodeBs), and NR NodeBs (gNBs or GNodeBs)). Other examples of nodes include, but are not limited to, core network functions, such as core network functions in a fifth-generation core network (5GC) (e.g., Access and Mobility Management Functions (AMF), Session Management Functions (SMF), and Network Slice Selection Functions (NSSF)).
[0019] Node 100 may be configured or operated to perform the methods and functions described herein, such as method 200 or 300 described below. Node 100 may include a processor (e.g., processing circuitry or logic) 102. It will be understood that node 100 may include one or more virtual machines running different software and / or processes. Therefore, node 100 may include one or more servers, switches, and / or storage devices, and / or may include cloud computing infrastructure running software and / or processes or infrastructure configured to perform in a distributed manner.
[0020] Processor 102 can control the operation of node 100 in the manner described herein. Processor 102 may include one or more processors, processing units, multi-core processors, or modules configured or programmed to control node 100 in the manner described herein. In certain embodiments, processor 102 may include multiple software and / or hardware modules, each configured to perform or be used to perform one or more steps of the functionality of node 100 as described herein.
[0021] Node 100 may include memory 104. In some embodiments, memory 104 of node 100 may be configured to store program code or instructions executable by processor 102 of node 100 to perform the functions described herein. Alternatively or additionally, memory 104 of node 100 may be configured to store any requests, resources, information, data, signals, etc., described herein. Processor 102 of node 100 may be configured to control memory 104 of node 100 to store any requests, resources, information, data, signals, etc., described herein.
[0022] It should be understood that node 100 can include as Figure 1 Other components that complement or are alternatives to the components shown. For example, in some embodiments, node 100 may include a communication interface. The communication interface can be used to communicate with other nodes in a communication network (e.g., other physical or virtual nodes). For example, the communication interface may be configured to send and / or receive requests, resources, information, data, signals, etc., to and / or from other nodes or network functions. The processor 102 of node 100 may be configured to control such a communication interface to send and / or receive requests, resources, information, data, signals, etc., to and / or from other nodes or network functions.
[0023] Node 100 is used to schedule radio resources to a user equipment (UE) in order to provide services to the UE's users. In short, in one embodiment, node 100 may be configured to acquire one or more physiological parameters of the UE's users and schedule resources to the UE based on those parameters to provide services to the UE with a predetermined quality of experience perceived by the user.
[0024] Figure 2 A method 200 for scheduling radio to a user equipment (UE) in a node to provide service to a user of the UE, according to some embodiments herein, is illustrated. In a first step, method 200 includes obtaining 202 one or more physiological parameters of the user of the UE. In a second step, the method includes scheduling 204 resources to the UE based on the one or more physiological parameters to provide the UE with a predetermined quality of experience perceived by the user.
[0025] The brain-aware radio resource management scheme described herein can lead to enhanced radio resource utilization. For example, as described in more detail below, in some embodiments, more resource blocks and / or higher transmission power can be allocated to users with a high latency-aware threshold compared to users with a low latency-aware threshold, while taking into account the latency requirements of the relevant application. Users with a high latency-aware threshold can correspond to elderly users, users performing activities, or users in a fatigued mental state. This can result in: i) lower costs and higher revenue for operators; ii) power savings in the network while maintaining the perceived QoS of users; iii) minimizing waste of radio resources and providing services to users more accurately based on their actual brain processing capabilities; and iv) energy efficiency.
[0026] Back Figure 2 More specifically, method 200 can be executed by node 100 as described above. In some embodiments, steps 202 and 204 of the method can be executed by a first processing module and a second processing module of processor 102 of node 100.
[0027] The node can schedule radio resources for a user equipment to provide services to the user equipment. More specifically, the UE may include a device capable of, configured to, arranged to, and / or operable to wirelessly communicate with network nodes and / or other wireless devices. Unless otherwise stated, the term "UE" may be used interchangeably with "wireless device (WD)" herein. Wireless communication may involve sending and / or receiving wireless signals using electromagnetic waves, radio waves, infrared waves, and / or other types of signals suitable for transmitting information through the air. Examples of UEs include, but are not limited to, smartphones, mobile phones, cellular phones, Voice over IP (VoIP) phones, wireless local loop phones, desktop computers, personal digital assistants (PDAs), wireless cameras, game consoles or devices, virtual reality devices or virtual reality consoles, music storage devices, playback devices, wearable terminal devices, wireless endpoints, mobile stations, tablet computers, laptop computers, devices embedded in laptop computers (LEE), devices mounted on laptop computers (LME), smart devices, wireless customer premises equipment (CPE), and personal wearable devices (e.g., watches, fitness trackers, etc.).
[0028] The service may include any application running on the user's device. In some embodiments, the service may need to interact with the user (e.g., in real time). For example, the service may include a game application, a virtual reality application, or any other user-centric application. In other embodiments, the service may include a video streaming application, a music streaming application, or any other application that sends audio or visual content to the user.
[0029] Node 100 can schedule resources to the UE. Node 100 can provide the scheduled resources to the UE (e.g., a node can schedule its own resources), or a node can schedule resources from another node so that the other node can provide services to the UE.
[0030] In step 202, method 200 may include acquiring one or more physiological parameters of the user of the UE (e.g., a user of a UE that has requested service). Physiological parameters in this sense may include, for example, any one or more of the following: heart rate, blood pressure, a measure of stress experienced by the user, and / or a measure of the user's activity level. For example, the activity level may be determined based on the identified type of activity (e.g., walking, dancing, standing still, running, cycling, etc.). Physiological parameters may also include a measure of user fatigue. However, those skilled in the art will understand that these are merely examples and other physiological parameters may also be acquired in step 202.
[0031] Typically, step 202, which involves acquiring one or more physiological parameters of a user of the UE, may include acquiring one or more physiological parameters from one or more sensors on the UE. Therefore, the UE may include one or more sensors that can be used to measure physiological parameters. For example, the UE may include any one or more of the following sensors: a sensor for measuring heart rate, a sensor for measuring blood pressure, a pulse oximeter (SpO2) sensor, a skin conductivity sensor, or any other sensor for measuring physiological parameters.
[0032] In some embodiments, the UE or node 100 may be configured to interact with another UE or device to acquire physiological measurements. For example, the UE or node 100 may interact with a user's health tracker or smartwatch to obtain physiological parameters.
[0033] These physiological parameters are most closely related to (e.g., mutually correlated with) a user's alertness and therefore their cognitive abilities at any given time. Therefore, allocating resources based on physiological parameters could provide an increasingly accurate method for measuring and allocating resources accordingly.
[0034] In some embodiments, step 202 may further include acquiring other human-centered parameters besides the physiological parameters described above. For example, the user's gender or age, or the time of day the user wishes to access the service. Typically, any parameters that are correlated with the user's cognitive speed may be further incorporated.
[0035] Note that users may consent to the acquisition or use of this information in this manner. In some embodiments, users may provide additional information, such as age or gender.
[0036] In step 204, the method includes scheduling resources to the UE based on one or more physiological parameters in order to provide the UE with a predetermined quality of experience perceived by the user.
[0037] In this sense, resources may include, for example, power and / or physical resource blocks (PRBs) allocated to provide services.
[0038] Typically, step 204 may include: if physiological parameters indicate that the user is alert or has a high perceptual speed (compared to physiological parameters indicating that the user is fatigued or has a low perceptual speed), then allocating more resources (e.g., using more resources to provide services).
[0039] In some embodiments, the user-perceived predetermined experience quality is based on whether the user perceives (or is able to perceive) latency in the service. In such embodiments, step 204 of scheduling resources to the UE may include determining a latency perception threshold based on one or more physiological parameters, and scheduling resources to the UE based on latency perception.
[0040] In this context, the latency-perceived threshold includes the amount of latency that a user can perceive. The more tired or fatigued a user is, the higher their latency-perceived threshold will be. This means they will tolerate higher latency, and the network will have a greater margin to accommodate that latency, thus allowing for a wider range of power allocation or PRB actions that the network can choose to minimize energy consumption.
[0041] Therefore, in some embodiments, the step of scheduling resources to the UE based on latency awareness may include: scheduling resources to provide services with a latency less than the latency awareness threshold.
[0042] The delayed perception threshold can be determined in various ways based on one or more physiological parameters, such as using a lookup table or a mapping between physiological parameters and the delayed perception threshold. Such a lookup table or mapping can be determined experimentally.
[0043] In some embodiments, a delay perception threshold can be determined based on one or more physiological parameters by using a first machine learning model to predict the delay perception threshold. The first machine learning model takes the one or more physiological parameters as input and outputs a prediction of the user's delay perception threshold based on the one or more physiological parameters.
[0044] Those skilled in the art will be familiar with machine learning models (e.g., models trained using machine learning processes). In short, machine learning can be used to find a prediction function for a given dataset; the dataset is typically a mapping between a given input and its output. The prediction function (or mapping function) is generated during a training phase, which involves providing the model with example inputs and corresponding benchmark true values (e.g., correct) of the output. The testing or validation phase then involves predicting the output for a given, previously unseen input. Applications of machine learning include, for example, curve fitting, face recognition, and spam filtering.
[0045] In this paper, the first machine learning model may include a supervised learning model, such as a classification or regression model. For example, in some embodiments, the machine learning model may include a neural network model, a random forest model, or a support vector regression model. While these are provided as example machine learning models, it will be understood that the teachings of this paper are more generally applicable to any type of model that can be trained to take one or more physiological parameters as input and output a prediction of a user's latency-perceived threshold.
[0046] As an example, a first machine learning model may include a (deep) neural network. Those skilled in the art will be familiar with neural networks, but in short, a neural network is a type of machine learning model that can be trained to predict a desired output for a given input data. The neural network is trained using training data, which includes example input data and the expected corresponding "correct" or benchmark true values. The neural network comprises multiple layers of neurons, each neuron representing a mathematical operation applied to the input data. The output of each layer in the neural network is fed into the next layer to produce an output. For each piece of training data provided to the neural network, the weights associated with the neurons are adjusted (e.g., using methods such as backpropagation and gradient descent) until optimal weights are found that produce predictions for the training examples that reflect the corresponding benchmark true values.
[0047] The first machine learning model may be trained using training data that includes training examples, wherein each training example includes: a set of example values of one or more physiological parameters of the example user, and a baseline real value delay-perceived threshold for the example user when obtaining example values of one or more physiological parameters of the example user.
[0048] The latency-perceived threshold can be determined on a per-user basis, for example, by asking users to indicate whether they perceived latency in the service when provided with different resource levels. In other words, the baseline latency-perceived threshold can be based on feedback from sample users regarding the quality of experience with sample services provided to them.
[0049] In some embodiments, a method is provided for training a supervised machine learning model to predict a user's latency-perceived threshold based on one or more physiological parameters of the user. The method includes providing training data to the machine learning model, the training data including training examples, each training example including: i) one or more physiological parameters of an example user, and ii) a latency-perceived threshold for the example user.
[0050] The following provides a detailed example of using a first machine learning model to predict a user's latency-aware threshold.
[0051] Training data collection process To train the first machine learning model, training data can be collected from multiple users under different states and mental conditions by asking them to rate the quality of the video as latency and packet loss increase in the system. For quality rating, metrics such as video distortion level, latency, and bitrate can be considered.
[0052] As mentioned above, the input features of the first machine learning model may include physiological parameters, including but not limited to:
[0053] Heart rate
[0054] - Activity Level: Steps / Second
[0055] - Pressure level
[0056] -Activity types: walking, dancing, stillness, running, cycling...
[0057] - Fatigue level
[0058] Other human-centered parameters can also be provided as input, such as:
[0059] -gender
[0060] -age
[0061] -Time of day
[0062] As mentioned above, this data can be collected via various sensors (such as sensors on the UE) or the user's associated devices (smartwatches / fitness trackers). The user is requested to notify the network of their satisfaction with the perceived signal quality.
[0063] Training program As described above, a first machine learning model can be trained using supervised machine learning techniques to learn the mapping between the model's input features and the desired output (which can be defined as the brain's delay perception threshold). Machine learning algorithms such as random forests and feedforward neural networks can be considered. The loss function can be defined as the mean squared error of all training examples. In this way, the first machine learning model can be used to predict the brain's delay perception threshold based on the acquired user physiological parameters.
[0064] Return to Figure 2 Once the user's perceived latency threshold is determined, method 200 may include scheduling resources to provide a service with a latency less than the perceived latency threshold. In this way, sufficient resources can be scheduled to the user so that the user perceives a high quality of service (e.g., no latency), but without over-supplying the user (the user would not be aware of the benefits of over-supplying resources due to their cognitive state).
[0065] Typically, the steps of scheduling resources to a UE may include using fewer resources to provide services if the user's perceived latency threshold is high (compared to a scenario where the user's perceived latency threshold is low). For example, the method may include providing services with fewer resources by transmitting service-related packets at lower power and / or allocating fewer resource blocks to the service. For instance, in cases where the user's perceived latency threshold is high (e.g., the user is fatigued or active), Node 100 may transmit services to the user at lower power and / or allocate fewer resource blocks. This allows for improvements in power savings, bandwidth allocation, and increased QoS by considering human characteristics along with radio parameters during the resource allocation process in this way. It also allows for the release of resources and their availability for other applications.
[0066] In some embodiments, step 204 of scheduling resources to a user equipment to provide services may include: scheduling resources to the user equipment using a reinforcement learning agent of a second machine learning model.
[0067] Those skilled in the art will be familiar with reinforcement learning and reinforcement learning agents; however, in short, reinforcement learning is a type of machine learning process in which a reinforcement learning agent (e.g., an algorithm) performs actions on a system to adjust the system according to a goal (which may include, for example, moving the system toward its optimal or preferred state). The reinforcement learning agent receives a reward based on whether each action alters the system in accordance with the goal (e.g., toward a preferred state) or contrary to the goal (e.g., away from a preferred state). Therefore, the reinforcement learning agent adjusts the parameters in the system with the goal of maximizing the received reward.
[0068] More formally, a reinforcement learning agent receives observations from the environment in state S and selects an action that maximizes the expected future reward r. Based on the expected future reward, the value function V for each state can be computed, and the optimal policy π that maximizes the long-term value function can be derived.
[0069] In the context of this disclosure, the telecommunications network is the “environment” in state S. “Observations” include physiological parameters and other human-related and / or radio-related features. Each “action” performed by the reinforcement learning agent includes a radio resource scheduling decision, which comprises a set of radio resource allocation parameters. Typically, the reinforcement learning agent in this paper receives feedback in the form of a reward or credit assignment each time it performs an adjustment (e.g., an action). As mentioned above, the goal of the reinforcement learning agent in this paper is to maximize the received reward.
[0070] Examples of reinforcement learning agents and reinforcement learning schemes that can be used for second machine learning models include, but are not limited to, Q-learning models, deep deterministic policy gradient (DDPG), deep Q-learning (DQN), and state-action-reward-state-action (SARSA).
[0071] In some embodiments, the reinforcement learning agent receives a positive reward if the scheduled resources meet the following conditions:
[0072] Delay < Delay-aware threshold [1]
[0073] For example, if the latency is below a threshold that is perceptible (or predicted to be perceptible) to the user in their given cognitive state.
[0074] As mentioned above, generally, the larger the latency-aware threshold, the greater the margin given to the network to meet that latency, and therefore the wider the range of power allocation or PRB actions the network can choose to minimize energy consumption.
[0075] In some embodiments, a reinforcement learning agent may further receive a positive reward if the scheduled resources maximize the following expression:
[0076] a*bitrate - b*energy [2]
[0077] Here, the parameter bitrate includes the bit rate at which the service is provided to the user, the parameter energy includes a measure of the energy required for the node to provide the service to the user at that bit rate, and a and b include weighted values.
[0078] a and b can include multi-objective weights (0) that can be used to achieve a trade-off between bit rate and energy efficiency. <a<1,0<b<1)。
[0079] In some embodiments, the reinforcement learning agent may further receive a positive reward if the scheduled resources meet the following conditions:
[0080] Delay <= network_delay_threshold [3]
[0081] The parameter `network_delay_threshold` includes parameters related to the network's allowed latency. For example, `network_delay_threshold` can be a parameter related to the requirements of the relevant application or the device requirements of the machine type. For instance, as an example of `network_delay_threshold`, for Ultra Reliable Low Latency Communication (URLLC) in 5G NR, a success probability of transmitting a 32-byte packet within 1ms is required to be 99.999%. In this way, the latency requirements of the relevant application can still be taken into account.
[0082] Constraint [1] encourages reinforcement learning agents to schedule sufficient resources to ensure that latency is less than the latency-aware threshold, thereby allocating resources more efficiently. If the latency drops to a level significantly below the latency-aware threshold, human users will not be able to discern the difference (compared to the case where the latency is just below the latency-aware threshold). This latency-aware threshold is determined based on physiological parameters through the human brain's ability. Constraints [2] and [3] guarantee that the latency is limited by the requirements of the relevant application and that the resources provided maximize the bit rate while minimizing energy usage. Each user or machine type device may have different applications with different QoS requirements. The variables in this optimization problem may be, but are not limited to, parameters such as: power allocation level, resource block allocation, and beam selection.
[0083] Method 200 may then include allocating resources to the UE according to the determined resource schedule. In other words, providing services to the UE using the scheduled resources.
[0084] The key difference between the proposed problem formulation and the traditional RB allocation problem lies in the QoS latency requirement, where the network explicitly considers the latency needs of the human brain. By taking into account the characteristics of the human brain, the network can avoid the resource waste caused by allocating more power to the UE solely based on applied QoS without considering how the human brain carrying the UE perceives QoS. The use of physiological parameters may be particularly beneficial because, for example, these parameters can be more closely correlated with the individual user's perception and alertness compared to other human-centered parameters.
[0085] In other embodiments, the reinforcement learning agent may receive rewards based on packet loss, for example, receiving a positive reward if the packet loss is below a threshold (e.g., a threshold required for the corresponding service).
[0086] The following provides a detailed implementation of a second machine learning model that includes a reinforcement learning agent.
[0087] Model initialization In this embodiment, the weights of the reinforcement learning agent can be initialized offline first, based on a conventional rule-based algorithm that only considers radio-related parameters. For example, for resource block allocation, round-robin or proportional fairness algorithms can be considered. The machine learning model can be, for example, a random forest, a convolutional neural network, or a feedforward neural network.
[0088] Model training The model is then trained to consider both channel state information and human-based metrics. During the training phase, users, user activity levels, and user states under different network conditions are considered. The state definition, action state, and reward function of the proposed reinforcement learning technique are summarized below:
[0089] State: Corresponds to a set of radio characteristics and human state characteristics:
[0090] Physiological parameters may include, but are not limited to:
[0091] o Heart rate
[0092] o Activity level: steps / second
[0093] o Pressure level
[0094] o Activity types: walking, dancing, stillness, running, cycling...
[0095] level of fatigue
[0096] Other human-centered parameters include:
[0097] o Gender
[0098] o age
[0099] o Time of day
[0100] o Radio-related characteristics:
[0101] Channel State Information (CSI)
[0102] RSRP / RSRQ / RSSI
[0103] For example, a high heart rate can reflect a user's state, such as stress levels and high activity levels, which is crucial for how the human brain perceives its environment, especially in video streaming applications. In this case, the brain is fatigued, so its latency perception is high, and the user may not be able to perceive the difference between very good and bad video (thus, for example, operators can take advantage of this to allocate less power / bandwidth). In a paper titled "Interactions between cardiac activity and conscious somatosensory perception" by Paweł Motyka, Martin Grund, Norman Forschack, Esra Al, Arno Villringer, and Michael Gaebler, the authors investigated the link between conscious perception and cardiac signals. They showed that the body's physiological state affects how we perceive the world. It is also noteworthy that stress levels, fatigue levels, and type of activity are also strongly correlated with heart rate (heart rates are higher under stress or strenuous physical activity), which therefore affects people's conscious perception and video quality assessment under different conditions.
[0104] In their paper, “Evaluating the Role of Content in Subjective Video Quality Assessment,” by Milan Mirkovic, Petar Vrgovic, Dubravko Culibrk, Darko Stefanovic, and Andras Anderla, the authors analyzed differences in human cognitive, emotional, and intentional responses to a set of videos commonly used for video quality assessment and to a specially selected set of videos containing content that could influence assessors’ judgments when discussing perceived video quality. They showed that cognitive mental activity was largely observed as a “rational” or “objective” calm state of mind. These activities are thought to be responsible for processing information acquired by people from their sensory systems through attention and memory. Furthermore, the paper emphasizes that other factors influencing subjective perceptions of video quality, therefore, can be considered when discussing video quality assessment tasks, and these factors relate to different population groups (by gender, culture, and demographics).
[0105] Therefore, the aforementioned human-related characteristics are important for resource allocation in wireless networks, where operators can consider quality of experience or people's perception of the service and adjust radio resources accordingly to achieve energy-efficient network management.
[0106] Action: A set of radio resource allocation parameters. Depending on the user-side application, one or more of the following radio-related parameters may be adjusted. Examples may include, but are not limited to:
[0107] o Transmit power level
[0108] The number of resource blocks allocated
[0109] beam selection
[0110] - Reward: The multi-objective weighted function defined in the optimization problem above, considering:
[0111] o rate
[0112] o delay
[0113] o Energy efficiency
[0114] The following schemes can be used to reward reinforcement learning agents:
[0115] Minimize - a*Bitrate + b*Energy [a]
[0116] Limited by: latency <= latency perception threshold[b]
[0117] Delay <= Machine type equipment requirements [c]
[0118] The delay-aware threshold in (3) can be inferred from the user rating results of the supervised learning scheme described above for the first machine learning model.
[0119] Actual operation In practice, whenever a given user registers on the network or uses sensors on the user's mobile device, node 100 may collect physiological parameters and / or other human-related data. This may be subject to, for example, the user's consent to the collection of data in this manner for use and / or local laws regarding the use of personal data.
[0120] If physiological parameters (and / or other human-related characteristics) are unavailable, radio resources can be allocated based solely on radio-related characteristics, as is done in traditional wireless networks. For example, node 100 can revert to a traditional scheduler.
[0121] Reinforcement learning agents (e.g., second models) can be continuously trained online based on user input that can be fed into the network (human-in-the-loop). For example, the network can request feedback from the user regarding the corresponding quality of the user's experience (e.g., a metric for assessing user-perceived latency). Safe exploration techniques can be employed to guarantee the QoS of end users when actions are randomly selected. Those skilled in the art will be familiar with safe exploration methods using reinforcement learning agents in communication networks. For example, research on network conditions that allow safe exploration includes the following paper: “Safe Exploration Algorithms for Reinforcement Learning Controllers”, by T. Mannuci, E. Kampen, C. Visser, and Q. Cu, IEEE Transactions on Neural Networks and Learning Systems, Vol. 29, No. 4, April 2018.
[0122] Now go to Figure 3 This illustrates an example of resource allocation. Figure 3In one embodiment, a method may include offline initialization of a reinforcement learning procedure 302 using only radio parameters. The network state 304 can then be observed. If network conditions allow for safe exploration, a random action (e.g., an "exploration action") can be taken on the current state. This action may include exploration actions regarding power level, beam selection, physical resource block (PRB) allocation, etc. A reward can then be assigned to this action in 310. If network conditions do not allow for safe exploration in box 306, it is determined in box 312 whether the user's physiological parameters are available. If available, the method includes taking an action known to have the highest reward value on the current state, considering both radio and human-related characteristics. This action may include power level selection, beam selection, or PRB allocation, as in box 308. If no physiological parameters are available in box 312, the method may include taking an action 316 with the highest reward value on the current state, considering only radio-related characteristics. Again, this action may include selecting a power level, beam selection, and / or PRB allocation. After step 314 or 316, the method includes assigning a reward 310 based on a reward function. Then, the method returns to step 304, preparing to take the next action.
[0123] In this way, exploration / development strategies for a second machine learning model can be designed that take into account the availability of physiological parameters when determining appropriate actions. This allows for safe exploration of online model updates.
[0124] In another embodiment, a computer program product including a computer-readable medium having computer-readable code contained therein, the computer-readable code being configured to cause the computer or processor, when executed by a suitable computer or processor, to perform one or more methods described herein.
[0125] Therefore, it will be understood that this disclosure also applies to computer programs adapted as practical embodiments, specifically computer programs on or in a carrier. Such programs may be in the form of source code, object code, intermediate source code, and object code (e.g., in partially compiled form), or any other form suitable for use in the implementation of methods according to the embodiments described herein.
[0126] It will also be understood that such a program can have many different architectural designs. For example, the program code implementing the functionality of the method or system can be subdivided into one or more subroutines. Many different ways of distributing functionality among these subroutines will be apparent to those skilled in the art. Subroutines can be stored together in an executable file to form a self-contained program. Such an executable file can include computer-executable instructions, such as processor instructions and / or interpreter instructions (e.g., Java interpreter instructions). Alternatively, one or more subroutines can be stored in at least one external library file and linked to the main program, for example, statically or dynamically at runtime. The main program contains at least one call to at least one subroutine. Subroutines can also include function calls to each other.
[0127] The carrier of a computer program can be any entity or device capable of carrying the program. For example, the carrier can include data storage, such as ROM (e.g., CD ROM or semiconductor ROM) or magnetic recording media (e.g., hard disk). Furthermore, the carrier can be a transmissible medium such as electrical or optical signals, which can be transmitted via cables or optical fibers or by radio or other means. When a program is contained within such a signal, the carrier can be constituted by such a cable or other device or apparatus. Alternatively, the carrier can be an integrated circuit with a program embedded within it, adapted to perform or be used in the execution of the relevant method.
[0128] By studying the accompanying drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments in practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. A single processor or other unit can implement the functions of several items as described in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not imply that combinations of these measures cannot be advantageously used. Computer programs can be stored / distributed on suitable media, such as optical storage media or solid-state media provided with or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems. Any reference numerals in the claims should not be construed as limiting the scope.
Claims
1. A computer-implemented method (200) in a node of a communication network for scheduling radio resources to a user equipment (UE) to provide services to a user of the UE, the method comprising: Obtain (202) one or more physiological parameters of the user of the UE; as well as Based on one or more physiological parameters, resources are scheduled (204) to the UE in order to provide the UE with a service having the predetermined experience quality perceived by the user. Wherein, the predetermined experience quality is based on whether the user perceives latency in the service, and wherein the step of scheduling (204) resources to the UE includes: Determine the latency perception threshold that the user can perceive based on one or more physiological parameters; and Resources are scheduled to the UE based on a latency-aware threshold.
2. The method according to claim 1, wherein, The step of scheduling resources to the UE based on the latency-aware threshold includes: scheduling the resources to provide services with a latency less than the latency-aware threshold.
3. The method according to claim 1 or 2, wherein, The step of determining the delayed perception threshold based on one or more physiological parameters includes: The delay perception threshold is predicted using a first machine learning model, which takes one or more physiological parameters as input and outputs a prediction of the delay perception threshold for the user based on the one or more physiological parameters.
4. The method according to claim 3, wherein, The first machine learning model is trained using training data that includes training examples, wherein each training example includes: a set of example values of one or more physiological parameters of the example user and a baseline real value delay-perceived threshold of the example user when the example values of the one or more physiological parameters are obtained.
5. The method according to claim 4, wherein, The baseline real-value latency perception threshold is based on feedback from the example user regarding the quality of experience of the example service provided to the example user.
6. The method according to any one of claims 1 to 2, wherein, The step of scheduling resources to the UE based on the latency-aware threshold includes: Compared to the case where the user's perceived latency threshold is low, when the user's perceived latency threshold is high, fewer resources are used to provide the service.
7. The method according to claim 6, wherein, Providing the service using fewer resources includes: Send packets related to the service at lower power; and / or A smaller number of resource blocks are allocated to the service.
8. The method according to any one of claims 1 to 2, wherein, The step of scheduling (204) resources to the user equipment to provide the service includes: A reinforcement learning agent using a second machine learning model schedules resources to the user equipment.
9. The method according to claim 8, wherein, The reinforcement learning agent receives a positive reward if the scheduled resources meet the following conditions: Delay < Delay perception threshold.
10. The method according to claim 8, wherein, The reinforcement learning agent receives a positive reward if the scheduled resources maximize the following expression: a*bitrate - b*energy; Wherein, the parameter bitrate includes the bit rate at which the service is provided to the user, the parameter energy includes a measure of the energy required by the node to provide the service to the user at the bit rate, and a and b include weighted values.
11. The method according to claim 8, wherein, The reinforcement learning agent receives a positive reward if the scheduled resources meet the following conditions: Delay <= network_delay_threshold; The parameter network_delay_threshold includes parameters related to the network's allowed latency.
12. The method according to any one of claims 1 to 2, wherein, The step of obtaining one or more physiological parameters of the user of the UE (202) includes: The one or more physiological parameters are acquired from one or more sensors on the UE.
13. The method according to any one of claims 1 to 2, wherein, The services include virtual reality applications or gaming applications.
14. The method according to any one of claims 1 to 2, wherein, The one or more physiological parameters include one or more of the following parameters: Heart rate; blood pressure; Pressure measurement; and The measure of a user's activity level.
15. The method according to any one of claims 1 to 2, wherein, The method also includes the following method steps: Resources are allocated to the UE according to the determined resource schedule.
16. A node (100) in a communication network for scheduling radio resources to a user equipment (UE) to provide services to a user of the UE, the node comprising: Memory (104) includes instruction data representing an instruction set; as well as Processor (102), configured to communicate with the memory and execute the instruction set, wherein the instruction set, when executed by the processor, causes the processor to: Obtain one or more physiological parameters of the user from the UE; and Resources are scheduled to the UE based on one or more physiological parameters in order to provide the UE with a service that has the predetermined quality of experience perceived by the user. Wherein, the predetermined experience quality is based on whether the user perceives latency in the service, and wherein, causing the processor to schedule resources to the UE includes causing the processor to: Determine the latency perception threshold that the user can perceive based on one or more physiological parameters; and Resources are scheduled to the UE based on a latency-aware threshold.
17. A node (100) in a communication network according to claim 16, wherein the node is adapted to perform any one of the methods defined in claims 2 to 15.
18. A computer program product comprising a computer-readable medium having computer-readable code contained therein, the computer-readable code being configured to cause, when executed by a suitable computer or processor, the computer or processor to perform the method according to any one of claims 1 to 15.