Reinforced learning in real-time communications
By applying a reinforcement learning system in real-time communication and automatically adjusting audio and video transmission parameters, the quality optimization problem in real-time communication is solved and the effect of continuously optimizing user experience is achieved.
Patent Information
- Application Number
- CN202510219557.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-10
- Filing Date
- 2020-06-08
- Publication Date
- 2025-05-30
AI Technical Summary
Bandwidth estimation, congestion control and video quality optimization remain a challenge in real-time communications, and existing technologies are difficult to continuously update to cope with new application needs and network behavior, resulting in a degradation of the end-user experience.
Reinforcement learning systems and methods are adopted to connect with the sending and receiving computing devices through the agent, and automatically adjust real-time audio and video transmission parameters to optimize the user's perceived quality experience. The system includes a reinforcement learning model that uses the current state, actions, and rewards (i.e., user experience quality) to determine the expected value of the sum of future rewards, thereby adjusting control strategies to optimize transmission parameters.
It realizes continuous optimization of user-perceived quality experience in real-time communication, quickly respond to changes in network conditions and application requirements, and reduces the downgrade of real-time audio and video transmission.
Smart Images

Figure CN120075207A_ABST
Abstract
Description
[0001] This application is a divisional application of a Chinese patent application with an application date of June 8, 2020, an application number of 202080049824.1, and an invention title of "Reinforcement Learning in Real-Time Communication". The Chinese patent application with an application number of 202080049824.1 is a Chinese national phase patent application entered from an international application with an international application number of PCT / US2020 / 036541. The international application with an international application number of PCT / US2020 / 036541 claims the priority of a US application with an application number of 16 / 507,933 filed on July 10, 2019. Background Art
[0002] Due to the frequent changes in network conditions and application requirements, bandwidth estimation, congestion control, and video quality optimization for real-time communication (e.g., voice and video conferencing) remain a difficult problem. The delivery of real-time media with high quality and reliability (e.g., quality of experience for end-users) requires continuous updates to respond to new application requirements and network behavior. The process of continuous updates can be a slow process, resulting in a degradation of the end-user experience.
[0003] Aspects of the present disclosure are made in view of these and other general considerations. Additionally, although relatively specific problems and examples of solving these problems may be discussed herein, it should be understood that the examples should not be limited to solving the specific problems identified in the background art or elsewhere in the present disclosure. Summary of the Invention
[0004] The present disclosure generally relates to systems and methods for implementing reinforcement learning in real-time communication. Certain aspects of the present disclosure relate to reinforcement learning for optimizing the quality perceived by users in real-time audio and video communication. An agent interfaces with a sending computing device and a receiving computing device to automatically adjust real-time audio and video transmission parameters in response to changing network conditions and / or application requirements. The sending computing device transmits real-time audio and / or video data. The receiving computing device receives the real-time audio and video transmission from the sending device and determines the actual quality of experience (QoE) perceived by the user, which is provided to the agent as a reward. The agent incorporates a reinforcement learning model that includes a control policy and a state-action value function. The agent observes the current state of the sending computing device and determines an estimate of the expected value of the sum of future rewards based on the current state, the current action (e.g., the current adjustment or set of adjustments to the transmission parameters at the sending computing device), and the reward provided by the receiving computing device. Based on the goal of maximizing the expected value of the sum of future rewards, the agent adjusts the control policy. Adjustments in the control policy can change the actions applied to the real-time audio and / or video data.
[0005] One aspect of the present disclosure relates to methods, systems, and articles of manufacture for optimizing expected user-perceived QoE in real-time communication. This aspect includes determining the current state of a sending computing device and the current actions of the sending computing device; the current actions include a plurality of transmission parameters. This aspect also includes transmitting real-time communication from the sending computing device to a receiving computing device. The real-time communication includes one or both of real-time audio communication and real-time video communication. Additionally, a reward (e.g., a QoE metric) is determined at the receiving computing device based on one or more parameters of the transmitted real-time communication received at the receiving computing device. Based on the current state, the current actions, and the reward, an expected value of the sum of future rewards is determined, and at least one of the plurality of transmission parameters of the sending computing device is changed to maximize the expected value of the sum of future rewards.
[0006] One aspect of the present disclosure relates to methods, systems, and articles of manufacture for a reinforcement learning model for optimizing expected user-perceived QoE in real-time communication. This aspect includes determining the current state of a sender and providing the current state to an agent communicating with the sender. This aspect also includes determining the current actions of the sender; the current actions are known to the agent and include a plurality of transmission parameters. This aspect also includes transmitting real-time communication from the sender to a receiver. The real-time communication includes one or both of real-time audio transmission and real-time video transmission. This aspect also includes receiving, at the agent, a reward determined at the receiver. The reward is based on one or more parameters associated with the real-time communication received at the receiver. The agent determines an expected value of the sum of future rewards based on the current state, the current actions, and the reward, and directs a change in at least one of the plurality of transmission parameters to maximize the expected value of the sum of future rewards. Training can be performed in a simulated environment, an emulated environment, or a real network environment.
[0007] This "Summary" is provided to introduce a selection of concepts in a simplified form that are further described below in the "Detailed Description". This "Summary" is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of examples will be set forth in part in the description that follows, and in part will be obvious from the description, or may be learned by practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Non-limiting and non-exhaustive examples are described with reference to the following figures.
[0009] Figure 1 An environment in which reinforcement learning in real-time communication as disclosed herein can be practiced is shown.
[0010] Figures 2A - 2C Additional details of an environment in which reinforcement learning in real-time communication as disclosed herein can be practiced are shown.
[0011] Figure 3 Shows a simulation training environment for an enhanced environment that maximizes the quality of experience (QoE) perceived by users in real-time communication.
[0012] Figure 4 Shows a simulation training environment for reinforcement learning that maximizes the QoE perceived by users in real-time communication.
[0013] Figure 5 Shows a real network training environment for reinforcement learning that maximizes the QoE perceived by users in real-time communication.
[0014] Figure 6 Is a block diagram showing example physical components of a computing device that can be used to practice aspects of the present disclosure.
[0015] Figure 7A And 7B Is a simplified block diagram of a mobile computing device that can be used to practice aspects of the present disclosure.
[0016] Figure 8 Is a simplified block diagram of a distributed computing system in which aspects of the present disclosure can be practiced.
[0017] Figure 9 Shows a tablet computing device for performing one or more aspects of the present disclosure. Detailed Description
[0018] Aspects of the present disclosure are described more fully hereinafter with reference to the accompanying drawings which form a part hereof. The different aspects of the present disclosure can be implemented in many different forms and should not be construed as limited to the aspects set forth herein; rather, these aspects are provided so that this disclosure will be thorough and complete and will fully convey the scope of these aspects to those skilled in the art. The aspects can be practiced as a method, system, or device. Thus, the aspects can take the form of a hardware implementation, a fully software implementation, or an implementation combining software and hardware aspects. Accordingly, the following detailed description should not be taken in a limiting sense.
[0019] The present disclosure generally relates to systems and methods for implementing reinforcement learning in real-time communication. Certain aspects of the present disclosure relate to reinforcement learning for optimizing the user-perceived quality in real-time audio and video communication. An agent interfaces with a sending computing device and a receiving computing device to automatically adjust real-time audio and video transmission parameters in response to changing network conditions and / or application requirements. The sending computing device transmits real-time audio and / or video data. The receiving computing device receives the real-time audio and video transmission from the sending device and determines the actual user-perceived quality of experience (QoE), which is provided to the agent as a reward. The agent incorporates a reinforcement learning model that includes a control policy and a state-action value function. The agent observes the current state of the sending computing device and determines an estimate of the expected value of the sum of future rewards based on the current state, the current action (e.g., the current adjustment or set of adjustments made to the transmission parameters at the sending computing device), and the reward provided by the receiving computing device. Based on the goal of maximizing the expected value of the sum of future rewards, the agent adjusts the control policy. Adjustments in the control policy can change the actions applied to the real-time audio and / or video data.
[0020] Accordingly, the present disclosure provides several technical benefits, including but not limited to a reinforcement learning model that is continuously updated and that immediately responds to adjusting the real-time audio and video transmission parameters of the sending computing device based on the goal of maximizing the expected value of the sum of future rewards. The real-time audio and video transmission parameters are immediately adjusted in response to changing network conditions and / or application requirements. Degradation of the transmitted real-time audio and video streams is minimized, which may occur in a previously used manual coding responsive update process of data transmission parameters for countering degradation.
[0021] Reference Figure 1 , shows an environment 100 for practicing reinforcement learning in real-time communication. Environment 100 includes a network 102 through which a plurality of computing devices 104 communicate via various communication links 106. The term "real-time" refers to data processing in which the received data is processed by the computing device almost immediately, e.g., at a level of computing device responsiveness where the user perceives it directly enough or such that the computing device can keep up with some external process.
[0022] Network 102 is any type of wired and / or wireless network that can transmit, receive, and exchange data, voice, and video traffic. Examples of networks include local area networks (LANs) that interconnect endpoints within a single domain and wide area networks (WANs) that interconnect multiple LANs, as well as subnets, metropolitan area networks, storage area networks, personal area networks (PANs), wireless local area networks (WLANs), campus area networks (CANs), virtual private networks (VPNs), passive optical networks, and the like.
[0023] The computing device 104 includes an endpoint of the network 102. The computing device 104 may include one or more general-purpose or special-purpose computing devices. Such devices may include, for example, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microcontroller-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, cellular phones, personal digital assistants (PDAs), gaming devices, printers, appliances, media centers, automotive embedded or attached computing devices, other mobile devices, distributed computing environments including any of the above systems or devices, etc. Further details regarding computing devices are in Figures 6 - 9 described.
[0024] Communication between the computing devices 104 travels through the link 106. The link may include any type of guided or unguided transmission medium capable of transmitting data, voice, and / or video from one computing device 104 to another computing device 104. Guided media transmit signals along a physical path. Examples of guided media include twisted pair cables, coaxial cables, fiber optics, etc. Unguided media transmit signals without using physical means to define the path taken by the signal. Examples of unguided media include radio waves, microwaves, infrared waves, etc.
[0025] Figures 2A - 2B An environment 200 is shown which, for illustrative purposes, includes a single sending computing device 204S and a single receiving computing device 204R that communicate in real time via the link 206 through the network 202. Although the sending computing device 204S is shown as only including sending capabilities, it should be recognized that the sending computing device 204S may also operate as a receiving computing device. Similarly, the receiving computing device 204R may also operate as a sending computing device. Thus, two-way real-time communication may occur between the sending computing device 204S and the receiving computing device 204R. The environment 200 communicates with the agent 206 in real time to implement reinforcement learning based on the real-time communication of data, which may include voice data and video data. Reinforcement learning optimizes the expected user-perceived quality in real-time communication by maximizing the expected value of the sum of future rewards. The agent 206 may include an encoding or application residing on one or both of the sending computing device 204S and the receiving computing device 204R. The agent 206 may also include an encoding or application residing on a computing device different from the sending computing device 204S or the receiving computing device 204R, such as a server computing device, a cloud computing device, etc.
[0026] As shown in the figure, the sending computing device 204S includes a data capture module 210, a data encoder module 212, and a data sender module 214. The data capture module 210 captures status data that represents the currently observed status of the sending computing device 204S. In the context of real-time audio and video communication, the currently observed status can include observed sending parameters that affect the transmission of real-time audio data and real-time video data. Observed sending parameters can include, for example, resolution, bit rate, frame rate, the stream to be sent, codec (encoding / decoding), the physical environment of the user (e.g., dark / light level, background noise, motion, etc.), or any other parameter that may affect real-time data transmission. The data encoder module 212 of the sending computing device 204S converts the status data into a specified format for real-time transmission over the network 202. The data sender module 214 sends the formatted status data to the network 202 in real time.
[0027] The receiving computing device 240R includes a data receiver module 220, a data decoder module 222, and a QoE metric module 224. The data receiver module 220 receives the formatted status data from the network 202 in real time and outputs network statistics to the proxy 206. Examples of network statistics include loss, jitter, round-trip time (RTT) (also known as network latency), receive rate, packet size, packet type, receive timestamp, sender timestamp, burst length in packet loss, gap between packet losses, or any other network statistics that can be used to evaluate the quality of the received audio and video data. The data decoder module 222 performs an operation inverse to that of the data encoder module 212 and extracts the received status data from the formatted status data in real time.
[0028] The QoE metric module 224 determines one or more metrics in the quality of experience (QoE) metric based on the extracted status data. The QoE metric represents the user-perceived quality of the received status data determined by a QoE machine learning model, such as a deep neural network (DNN) or other suitable model. The QoE machine learning model analyzes various received parameters, such as the payloads of the received audio and video data streams, where the payload is the part of the received data that is the actual expected message. The analysis of the payloads of the audio and video streams may include using one or more predefined objective models that approximate the results of subjective quality assessments (e.g., ratings of quality by human observers). In some examples, the objective model may include one or more models for evaluating the quality of real-time audio (e.g., the Perceptual Evaluation of Audio Quality (PEAQ) model, the PEMO-Q model, the Signal-to-Noise Ratio (PSNR) model, or any other objective model that can evaluate the received real-time audio signal). In some examples, the objective model may include one or more models for evaluating the quality of real-time video (e.g., the Full Reference (FR) model, the Reduced Reference (RR) model, the No Reference (NR) model, the Peak Signal-to-Noise Ratio (PSNR) model, the Structural Similarity Index (SSIM) model, or any other model that can evaluate the received real-time video signal).
[0029] In some aspects, the QoE machine learning model may additionally analyze network statistics and statistics of the receiving computing device 204R as received parameters to determine one or more QoE metrics. As described herein, examples of network statistics include loss, jitter, round-trip time (RTT) (also known as network latency), receive rate, packet size, packet type, receive timestamp, sender timestamp, burst length in packet loss, gap between packet losses, or any other network statistics that can be used to evaluate the quality of the received audio and video data. Examples of statistics of the receiving computing device 204R include display size, display window size, device type, whether to use a hardware or software encoder / decoder, etc. In some aspects, the QoE machine learning model may additionally analyze user (e.g., human) feedback as a received parameter to determine one or more QoE metrics. User feedback may be provided, for example, through user ratings or surveys to indicate their personal quality of experience, such as their perception of the quality of the audio and video received at the receiving computing device 204R. The one or more QoE metrics determined to represent the user-perceived audio and / or video quality are transmitted to the agent 206.
[0030] Agent 206 includes a status module 230 and a reinforcement learning model 232. In some aspects, the reinforcement learning model 232 can incorporate any suitable reinforcement learning algorithm (a learning algorithm in which actions occur, results are observed, and the next action takes into account the results of the first action based on a reward signal). Reinforcement learning algorithms can include, for example, actor-critic, q-learning, policy gradient, temporal difference, Monte Carlo tree search, or any other reinforcement learning algorithm suitable for the data involved. The reinforcement learning model 232 actively controls the data transmission parameters sent to the computing device 204S in real time.
[0031] Figure 2B An example of an actor-critic reinforcement learning model 232 including a control policy 234 and a state-action value function 236 is shown. Figure 2C An example of an actor-critic architecture is provided. Actor-critic reinforcement learning is a temporal difference learning method in which the control policy 234 is independent of the estimated state value function 236, which in the current context is the expected value of the sum of future rewards. The control policy 234 includes an actor because it is used to select actions, such as the data transmission parameters of the sending computing device, and the state-action value function 236 is a critic because it evaluates the actions made by the control policy 234. The state-action value function 236 learns and evaluates the current control policy 234.
[0032] The control policy 234 includes a first machine learning model within the agent 206, such as a neural network, that produces one or more output actions in the form of one or more changes to one or more of the data transmission parameters used by the sending computing device 204S. The output actions are designed to optimize the expected user-perceived quality of experience (QoE) of audio and video data based on the maximization of the expected value of the sum of future rewards determined by the state-action value function 236. Examples of data transmission parameters include transmission rate, resolution, frame rate, object events provided to the quantization parameter (QP), forward error correction (FEC), or any other controllable parameter that can be used to modify the quality of the transmission of state data from the sending computing device 204S to the receiving computing device 204R.
[0033] The state-action value function 236 includes a second machine learning model within the agent 206, such as a neural network, and the value function of the second machine learning model is trained to predict or estimate the expected value of the future reward sum. Based on the current state of the sending computing device, the current action (e.g., the current transmission parameters for transmitting real-time audio and / or video data), and the reward provided by the receiving computing device, the expected value of the future reward sum is determined. The control policy adjusts the output action in response to the determination of the expected value. The control policy 234 can be trained together with the state-action value function 236 or can be obtained based on the already trained state-action value function 236.
[0034] In some aspects, during the training of the actor-critic reinforcement learning model 232 in Figures 2B - 2C , the agent 206 does not always need to follow the actions of the control policy 234. Instead, the agent 206 can explore other actions (e.g., other modifications to the data transmission parameters of the sending computing device 204S), which allows the agent 206 to improve the reinforcement learning model 232. The agent 206 can explore other actions through one or more exploration strategies (e.g., epsilon-greedy).
[0035] In some aspects, the control policy 234 of the reinforcement learning model 232 can be separated from its learning environment and deployed as a real-time model in a client (e.g., the sending computing device and / or the receiving computing device). The transfer to the real-time model can be achieved through one or more model transfer tools such as ONNX (Open Neural Network Exchange), tflite (TensorFlowLite), etc.
[0036] Refer to Figures 3 - 5 , one or more of the simulation environment 300, the emulation environment 400, and the real network environment 500 can be used to train the agent 206. Which environment to use depends on the requirements of data collection speed and data diversity. In the simulation environment of FIG. 300, all processes of the sending computing device 204S (including the processes of the data capture module 210, the data encoder module 212, and the data sender module 214), all processes of the receiving computing device 204R (including the processes of the data receiver module 220, the data decoder module 222, and the QoE metric module 224), and the network 202 are simulated. In Figure 4In the simulation environment 400, the sending computing device 204S is replicated in a first simulation including a simulated sending process 404S, the receiving computing device 204R is replicated in a second simulation including a simulated receiving process 404R, and the network 202 is replicated in a third simulation including a network simulation 402. In some aspects, the physical sending computing device and the physical receiving computing device can be used in conjunction with the simulated network. In the real network environment of FIG. 500, the physical sending computing device 204S, the physical receiving computing device 204R, and the physical network 202 are used.
[0037] Which environment to use to train the agent 206 depends on the data collection speed and the data collection diversity requirements. For example, network simulation tools such as ns-2 or ns-3 (which are discrete event network simulators) can be used in the simulation environment 300 for fast data collection and training. Network simulation tools such as NetEm (which is an enhancement of the Linux traffic control facility that allows adding latency, packet loss, duplication, and other characteristics of outgoing transmission packets from a selected network interface) can be used in the simulation environment 400 to allow real code to run in a controlled environment. Such a controlled environment allows testing of communication applications (e.g., Skype, Microsoft Teams, WhatsApp, WeChat, etc.) in an environment with reproducible network conditions. Using the real network (e.g., cellular, Wi-Fi, Ethernet, etc.) of a real Internet service provider (ISP) in the real network environment 500 provides the most realistic test environment and allows online learning of the conditions experienced by end users. In some aspects, the same reinforcement learning strategy can be used in the simulation, simulation, or real network environment, however, each environment will provide different performance. Alternatively or additionally, transfer learning can be used to train the agent 206, where manually encoded rules are used to train the agent 206, and the manually encoded rules were previously created in response to new application requirements and / or network behavior related to real-time audio and video data streaming.
[0038] Once trained, the agent 206 is applied in a live network environment for real-time audio and video communication. Within the live network, the reinforcement learning model 232 is continuously updated based on the transmission of real-time audio and video data streams from the sending computing device (e.g., device 204S) to the receiving computing device 204R. In some aspects, the sending computing device (e.g., device 204S) can include a single agent 206 or multiple agents 206, and the agents 206 operate to modify real-time audio and video data transmission parameters, with each agent modifying only one data transmission parameter or multiple data transmission parameters. In some aspects, the receiving computing device (e.g., device 204R) can determine one QoE or multiple QoEs. One or more QoEs can be provided to a single agent 206 or multiple agents 206.
[0039] Accordingly, based on the continuous real-time updates of the agent 206, the agent 206 and the sending computing device 204S are immediately (e.g., in real time) updated to continuously optimize the expected user-perceived quality in real-time audio and video communication by maximizing the expected value of the sum of future rewards, rather than suffering from degraded real-time audio and video transmission, which would otherwise result in a context of an environment where only manual encoding is used to respond to network condition changes and / or application demand changes.
[0040] Figures 6 - 9 And the associated description provides a discussion of various operating environments in which aspects of the present disclosure may be practiced. However, with regard to Figures 6 - 9 the devices and systems illustrated and discussed are for purposes of example and illustration and do not limit the numerous computing device configurations that may be used to practice aspects of the present disclosure, as described herein.
[0041] Figure 6 is a block diagram illustrating physical components (e.g., hardware) of a computing device 600 that may be used to practice aspects of the present disclosure. The computing device components described below may have computer-executable instructions for implementing reinforcement learning that maximizes the user-perceived QoE in real-time communication on a computing device (e.g., the sending computing device 204S and the receiving computing device 204R), including computer-executable instructions for a reinforcement learning application 620 that may be executed to implement the methods disclosed herein. In a basic configuration, the computing device 600 may include at least one processing unit 602 and a system memory 604. Depending on the configuration and type of the computing device, the system memory 604 may include, but is not limited to, volatile memory (e.g., random access memory), non-volatile memory (e.g., read-only memory), flash memory, or any combination of such memories. The system memory 604 may include an operating system 605 and one or more program modules 606, such as one or more components with respect to FIG. 2, particularly a data capture, data encoder, and data sender module 611 (e.g., a data capture module 210, a data encoder module 212, and a data sender module 214), a data receiver, data decoder, and QoE metric module 613 (e.g., a data receiver module 220, a data decoder module 222, and a QoE metric module 224), and / or an agent module 615 (e.g., the agent 206).
[0042] For example, the operating system 605 may be adapted to control the operation of the computing device 600. Additionally, embodiments of the present disclosure may be practiced in conjunction with a graphics library, other operating systems, or any other application presentation and are not limited to any particular application or system. This basic configuration is in Figure 6is shown by the components within the dashed line 608. The computing device 600 may have additional features or functionality. For example, the computing device 600 may also include additional data storage devices (removable and / or non-removable), such as, for example, magnetic disks, optical disks, or magnetic tapes. Such additional storage is shown in Figure 6 by removable storage device 609 and non-removable storage device 610. Any number of program modules and data files may be stored in the system memory 604. When executed on the processing unit 602, the program modules 606 (e.g., reinforcement learning application 620) may perform processes including, but not limited to, aspects described herein.
[0043] In addition, embodiments of the present disclosure may be practiced in a circuit including discrete electronic elements, a packaged or integrated electronic chip containing logic gates, a circuit utilizing a microprocessor, or on a single chip containing electronic elements or a microprocessor. For example, embodiments of the present disclosure may be practiced via a system-on-chip (SOC), where Figure 6 each or many of the components shown may be integrated onto a single integrated circuit. Such an SOC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functionalities, all of which are integrated (or "burned") onto a chip substrate as a single integrated circuit. When operating via an SOC, the functionality regarding the client switching protocol described herein may be operated via dedicated logic integrated on a single integrated circuit (chip) with other components of the computing device 600. Embodiments of the present disclosure may also be practiced using other technologies capable of performing logical operations such as, for example, AND, OR, and NOT, including but not limited to mechanical, optical, fluidic, and quantum technologies. In addition, embodiments of the present disclosure may be practiced within a general-purpose computer or in any other circuit or system.
[0044] The computing device 600 may also have one or more input devices 612, such as a keyboard, mouse, pen, voice or speech input device, touch or swipe input device, etc. One or more output devices 614 (such as a display, speaker, printer, etc.) may also be included. The foregoing devices are examples and other devices may be used. The computing device 600 may include one or more communication connections 616 that allow communication with other computing devices 650. Examples of suitable communication connections 616 include, but are not limited to, radio frequency (RF) transmitter, receiver, and / or transceiver circuitry; universal serial bus (USB), parallel, and / or serial ports.
[0045] The term computer-readable medium as used herein can include computer storage media. Computer storage media can include volatile and non-volatile removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, or program modules. System memory 604, removable storage device 609, and non-removable storage device 610 are all examples of computer storage media (e.g., memory storage devices). Computer storage media can include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article that can be used to store information and can be accessed by computing device 600. Any such computer storage media can be part of computing device 600. Computer storage media does not include carrier waves or other propagated or modulated data signals.
[0046] Communication media can be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information delivery media. The term "modulated data signal" can describe a signal in which one or more characteristics are set or changed in such a manner as to encode information in the signal. By way of example and not limitation, communication media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0047] Figure 7A and 7B Illustrated is a mobile computing device 700, e.g., a mobile phone, smartphone, wearable computer (such as a smartwatch), tablet computer, laptop computer, etc., and embodiments of the present disclosure can be practiced using these mobile computing devices 700. In some aspects, the client can be a mobile computing device. Referring to Figure 7A, which shows one aspect of a mobile computing device 700 for implementing these aspects. In a basic configuration, the mobile computing device 700 is a handheld computer with both input and output elements. The mobile computing device 700 typically includes a display 705 and one or more input buttons 710 that allow a user to enter information into the mobile computing device 700. The display 705 of the mobile computing device 700 can also be used as an input device (e.g., a touchscreen display). If included, the optional side input element 715 allows for further user input. The side input element 715 can be a rotary switch, a button, or any other type of manual input element. In alternative aspects, the mobile computing device 700 can incorporate more or fewer input elements. For example, in some embodiments, the display 705 may not be a touchscreen. In yet another alternative embodiment, the mobile computing device 700 is a portable telephone system, such as a cellular phone. The mobile computing device 700 may also include an optional keypad 735. The optional keypad 735 can be a physical keypad or a "soft" keypad generated on a touchscreen display. In various embodiments, the output elements include a display 705 for presenting a graphical user interface (GUI), a visual indicator 720 (e.g., a light-emitting diode), and / or an audio transducer 725 (e.g., a speaker). In some aspects, the mobile computing device 700 incorporates a vibration transducer for providing haptic feedback to the user. In yet another aspect, the mobile computing device 700 incorporates input and / or output ports, such as an audio input (e.g., a microphone jack), an audio output (e.g., a headphone jack), and a video output (e.g., an HDMI port), for sending signals to or receiving signals from external devices.
[0048] Figure 7B is a block diagram showing the architecture of one aspect of a mobile computing device. That is, the mobile computing device 700 can incorporate a system (e.g., an architecture) 702 to implement some aspects. In one embodiment, the system 702 is implemented as a "smartphone" capable of running one or more applications (e.g., a browser, email, calendar, contact manager, messaging client, game, and media client / player). In some aspects, the system 702 is integrated as a computing device, such as an integrated personal digital assistant (PDA) and a wireless phone.
[0049] One or more applications 766 may be loaded into the memory 762 and run on or in association with the operating system 764. Examples of applications include a telephone dialer, an email program, a personal information management (PIM) program, a word processing program, a spreadsheet program, an Internet browser program, a messaging program, and the like. The system 702 also includes a non-volatile storage area 768 within the memory 762. The non-volatile storage area 768 may be used to store persistent information that should not be lost when the system 702 is powered off. The applications 766 may use the information in the non-volatile storage area 768 and store information in the non-volatile storage area 768, such as emails or other messages used by an email application. A synchronization application (not shown) also resides on the system 702 and is programmed to interact with a corresponding synchronization application residing on a host to keep the information stored in the non-volatile storage area 768 synchronized with the corresponding information stored at the host. It should be understood that other applications may be loaded into the memory 762 and run on the mobile computing device 700, including instructions for providing a consensus determination application as described herein (e.g., a message parser, a suggestion interpreter, an opinion interpreter, and / or a consensus demonstrator, etc.).
[0050] The system 702 has a power supply 770, which may be implemented as one or more batteries. The power supply 770 may also include an external power source, such as an AC adapter or a power dock stand that supplements or recharges the battery.
[0051] The system 702 may also include a radio interface layer 772 that performs the functions of transmitting and receiving radio frequency communications. The radio interface layer 772 facilitates a wireless connection between the system 702 and the "outside world" via a communication carrier or service provider. Transmissions to and from the radio interface layer 772 are under the control of the operating system 764. In other words, communications received by the radio interface layer 772 may be propagated to the applications 766 via the operating system 764, and vice versa.
[0052] The visual indicator 720 may be used to provide visual notifications, and / or the audio interface 774 may be used to via the audio transducer 725 (e.g., Figure 7AThe illustrated audio transducer 725 generates an audible notification. In the illustrated embodiment, the visual indicator 720 is a light-emitting diode (LED) and the audio transducer 725 can be a speaker. These devices can be directly coupled to the power supply 770 such that when activated, they remain on for the duration specified by the notification mechanism even if the processor 760 and other components may be turned off to conserve battery power. The LED can be programmed to remain lit indefinitely until the user takes an action to indicate the powered-on state of the device. The audio interface 774 is used to provide audible signals to and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 725, the audio interface 774 can also be coupled to a microphone to receive audible input, such as to facilitate a telephone conversation. According to an embodiment of the present disclosure, the microphone can also be used as an audio sensor to facilitate control of the notification, as described below. The system 702 can also include a video interface 776 that enables the peripheral device 730 (e.g., an on-vehicle camera) to operate to record still images, video streams, etc. The audio interface 774, the video interface 776, and the keyboard 735 can be operated to generate one or more messages as described herein.
[0053] The mobile computing device 700 implementing the system 702 can have additional features or functionality. For example, the mobile computing device 700 can also include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or magnetic tapes. Such additional storage is illustrated by the non-volatile storage area 768 in Figure 7B .
[0054] Data / information generated or captured by the mobile computing device 700 and stored via the system 702 can be stored locally on the mobile computing device 700 as described above, or the data can be stored on any number of storage media that can be accessed by the device via the radio interface layer 772 or via a wired connection between the mobile computing device 700 and a separate computing device associated with the mobile computing device 700 (e.g., a server computer in a distributed computing network such as the Internet). It should be understood that such data / information can be accessed via the radio interface layer 772 or via the distributed computing network via the mobile computing device 700. Similarly, such data / information can be easily transmitted between computing devices for storage and use according to well-known data / information transmission and storage means, including email and collaborative data / information sharing systems.
[0055] It should be understood that Figure 7A and 7B are described for purposes of illustrating the method and system and are not intended to limit the present disclosure to a particular sequence of steps or a particular combination of hardware or software components.
[0056] Figure 8 FIG. 1 shows an aspect of the architecture of a system for processing data received at a computing system from a remote source, such as a general-purpose computing device 804 (e.g., a personal computer), a tablet computing device 806, or a mobile computing device 808 as described above. The content displayed at the server device 802 can be stored in different communication channels or other storage types. For example, a directory service 822, a web portal 824, a mailbox service 826, an instant messaging repository 828, or a social networking service 830 can be used to receive and / or store various messages. A reinforcement learning application 821 can be employed by a client communicating with the server device 802, and / or a reinforcement learning application 820 can be used by the server device 802. The server device 802 can provide data to and from client computing devices, such as a general-purpose computing device 804, a tablet computing device 806, and / or a mobile computing device 808 (e.g., a smart phone), via a network 815. For example, the computer system described above can be embodied in a general-purpose computing device 804 (e.g., a personal computer), a tablet computing device 806, and / or a mobile computing device 808 (e.g., a smart phone). In addition to receiving graphical data that can be pre-processed at a graphics initiation system or post-processed at a receiving computing system, any of these embodiments of the computing device can obtain content from a repository 816.
[0057] It should be understood that Figure 8 is described for purposes of illustrating the method and system and is not intended to limit the disclosure to a particular sequence of steps or a particular combination of hardware or software components.
[0058] Figure 9 FIG. 9 shows an exemplary tablet computing device 900 that can perform one or more aspects disclosed herein. Additionally, the aspects and functionality described herein can operate on a distributed system (e.g., a cloud-based computing system), where application functionality, memory, data storage and retrieval, and various processing functions can operate remotely from each other via a distributed computing network, such as the Internet or an intranet. Various types of user interfaces and information can be displayed via an on-board computing device display or via a remote display unit associated with one or more computing devices. For example, various types of user interfaces and information can be displayed on and interacted with a wall surface on which various types of user interfaces and information are projected. Interaction with the various computing systems that can be used to practice embodiments of the present invention includes keystroke entry, touchscreen entry, voice or other audio entry, gesture entry, where the associated computing device is equipped with detection (e.g., camera) functionality for capturing and interpreting user gestures to control the functionality of the computing device, etc.
[0059] It should be understood that Figure 9It is described for the purpose of illustrating the method and system and is not intended to limit the present disclosure to a particular sequence of steps or a particular combination of hardware or software components.
[0060] The description and illustration of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed present disclosure in any way. The aspects, examples, and details provided in this application are considered sufficient to convey ownership and enable others to make and use the best mode of the claimed disclosure. The claimed disclosure should not be construed as limited to any aspect, example, or detail provided in this application. The various features (structural and methodological), whether shown and described in combination or separately, are intended to be selectively included or omitted to produce embodiments having a particular set of features. Having provided the description and illustration of this application, those skilled in the art can envision variations, modifications, and alternative aspects that fall within the spirit of the broader aspects of the general inventive concept embodied in this application but do not depart from the broader scope of the claimed disclosure.
Claims
1. A method for optimizing the expected user-perceived quality of experience (QoE) in real-time communication between a sending computing device and a receiving computing device, comprising: determining, by the sending computing device, a current state of the sending computing device; determining a current action of the sending computing device; transmitting real-time communication from the sending computing device to the receiving computing device, where the real-time communication includes one or more of real-time audio communication and real-time video communication; receiving a reward and network statistics, where the reward is based on one or more reception parameters associated with the transmitted real-time communication received at the receiving computing device; determining an expected value of a sum of multiple future rewards based on the current state, the current action, the network statistics, and the reward; and changing at least one transmission parameter of the multiple transmission parameters of the sending computing device to maximize the expected value of the sum of the multiple future rewards.
2. The method according to claim 1, wherein a state-action value function of a reinforcement learning model determines the expected value of the sum of the multiple future rewards.
3. The method according to claim 2, further comprising providing an output of the state-action value function to a control policy learning model of the reinforcement learning model, and changing the at least one transmission parameter of the multiple transmission parameters by the control policy learning model based on the output of the state-action value function.
4. The method according to claim 1, wherein the reward includes a user-perceived QoE metric based on the one or more reception parameters associated with the transmitted real-time communication received at the receiving computing device.
5. The method according to claim 4, further comprising using a QoE machine learning model to determine the user-perceived QoE, where the QoE machine learning model evaluates the payload of the transmitted real-time communication received at the receiving computing device.
6. The method according to claim 4, using a QoE machine learning model to determine the user-perceived QoE, where the QoE machine learning model evaluates: the network statistics at the receiving computing device; receiving computing device statistics; and user feedback on the transmitted real-time communication received at the receiving computing device.
7. The method according to claim 1, wherein the at least one transmission parameter of the multiple transmission parameters comprises: a transmission rate parameter, a resolution parameter, a frame rate parameter, a quantization parameter QP, or a forward error correction (FEC) parameter.
8. The method according to claim 1, wherein the sending computing device additionally operates as a receiving computing device and wherein the receiving computing device additionally operates as a sending computing device for two-way real-time communication.
9. A method for training a reinforcement learning model for optimizing the expected user-perceived quality of experience (QoE) in real-time communication, the method comprising: The sending computing device determines the current state of the sending computing device; Determine the current action of the sending computing device; Transmit real-time communication from the sending computing device to a receiving computing device, where the real-time communication includes one or more of real-time audio communication and real-time video communication; Receive network statistics of a reward and a network, where the first computing device and the second computing device communicate through the network, and where the reward is based on one or more received parameters associated with the real-time communication received at the receiving computing device; Determine an expected value of a sum of multiple future rewards based on the current state, the current action, the network statistics, and the reward; And Change at least one of the multiple transmission parameters to maximize the expected value of the sum of the multiple future rewards.
10. The method according to claim 9, wherein the sending computing device, the receiving computing device, and the network are simulated.
11. The method according to claim 10, wherein the sending computing device, the receiving computing device, and the network are simulated using discrete events.
12. The method according to claim 10, wherein the current state of the sending computing device is determined based on the network statistics indicating the quality of the real-time communication received at the receiving computing device.
13. The method according to claim 9, wherein each of the sending computing device and the receiving computing device executes a communication application, and wherein one or more conditions of the network are controlled according to one or more predetermined parameters.
14. The method according to claim 9, wherein the network includes a live, real network.
15. The method according to claim 14, wherein the sending computing device, the receiving computing device, and the network are in a live environment, and the method further includes continuously training the agent based on live real-time communication transmissions.
16. A system for optimizing an expected user-perceived quality of experience QoE in real-time communication, Comprising: A memory storing executable instructions; And A processor executing the executable instructions, which when executed cause the processor to: The sending computing device determines the current state of the sending computing device; Determine the current action of the sending computing device; Transmit real-time communication from the sending computing device to a receiving computing device, where the real-time communication includes one or more of real-time audio communication and real-time video communication; Receive a reward and network statistics, where the reward is based on a parameter associated with the transmitted real-time communication received at the receiving computing device; Determine an expected value of a sum of multiple future rewards based on the current state, the network statistics, the current action, and the reward; And Change at least one of the multiple transmission parameters of the sending computing device to maximize the expected value of the sum of the multiple future rewards.
17. The system according to claim 16, further comprising operating the sending computing device as a receiving computing device for two-way real-time communication.
18. The system according to claim 16, wherein the at least one transmission parameter among the plurality of transmission parameters comprises: a transmission rate parameter, a resolution parameter, a frame rate parameter, a quantization parameter QP, or a forward error correction FEC parameter.
19. The system according to claim 16, further comprising performing the determination of the expected value of the sum of the plurality of future rewards by using a reinforcement learning model.
20. The system according to claim 19, wherein the reinforcement learning model comprises an actor / critic model, a q-learning model, a policy gradient model, a temporal difference model, or a Monte Carlo tree search model.