QoE-driven low-orbit satellite holographic video cross-layer adaptive transmission method

By employing a QoE-driven cross-layer adaptive transmission method for holographic video in low-Earth orbit satellites, combined with a 3D block model and deep reinforcement learning, the transmission bottleneck of holographic video in low-Earth orbit satellite systems is solved, achieving stable transmission of holographic video and maximizing long-term user experience quality.

CN121814920APending Publication Date: 2026-04-07CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing low-Earth orbit satellite systems cannot effectively meet the extreme data requirements and dynamic channel characteristics of holographic videos, resulting in transmission bottlenecks and a lack of cross-layer collaborative design to achieve strict latency constraints and maximize long-term user experience quality.

Method used

A QoE-driven cross-layer adaptive transmission method for low-orbit satellite holographic video is adopted. Through 3D block model, Rician fading channel modeling and deep reinforcement learning, combined with video quality level, content compression strategy and physical layer power allocation, cross-layer collaborative optimization is achieved. A hybrid algorithm design combining deep reinforcement learning and convex optimization is used to ensure stable transmission in dynamic environments.

Benefits of technology

In scenarios involving time-varying low-orbit satellite channels and user field-of-view drift, the end-to-end latency of holographic video is stabilized at the millisecond level, supporting stable transmission of immersive services such as holographic conferencing and virtual collaboration, and improving the long-term user experience quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121814920A_ABST
    Figure CN121814920A_ABST
Patent Text Reader

Abstract

The invention relates to a QoE-driven low-orbit satellite holographic video cross-layer adaptive transmission method, and belongs to the technical field of wireless communication. The method comprises the following steps: constructing a low earth orbit satellite end-to-end holographic video stream transmission system model, completing holographic object space division and user view field dynamic mapping through a 3D block model, realizing Rician fading channel modeling by means of a communication model, quantizing end-to-end time delay by means of a time delay model, and establishing a low earth orbit satellite end-to-end holographic video stream transmission system. Based on the QoE model, fusing the spatial importance and the time domain stability to measure the user experience; a joint decision variable of a video quality level, a content compression strategy and physical layer power distribution is designed to realize cross-layer collaboration, a real-time decision sequence is converted into a time domain coupled continuous strategy optimization problem, and long-term user experience quality maximization is taken as a target. End-to-end time delay, NOMA continuous interference elimination feasible region and satellite power upper limit are restrained at the same time; and solving by using a joint flow strategy of DRL learning quality and compression.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of wireless communication, and relates to a QoE-driven holographic video cross-layer adaptive transmission method for low-orbit satellites. BACKGROUND

[0002] The deployment of low earth orbit (LEO) mega-constellations is opening up new frontiers for global immersive applications, including holographic communications, and bringing unprecedented opportunities for application scenarios such as holographic conferencing and virtual collaboration. However, a fundamental mismatch threatens this vision: the dynamic and resource-constrained nature of LEO systems cannot meet the stringent millisecond-level delay requirements posed by multi-gigabit raw holographic video data streams per second. This mismatch creates a serious bottleneck, hindering the deployment of next-generation immersive services on satellite networks. In order to release its potential, a new paradigm is needed - that is, an intelligent cross-layer optimization method that can jointly design the streaming strategy of the application layer and the resource allocation of the physical layer to cope with the extreme data requirements and variable channel conditions inherent in the LEO environment.

[0003] Existing research, although extensive, has not yet been able to provide a unified solution to this particular cross-layer challenge. Previous research on adaptive streaming, although of foundational significance, mainly targets traditional video formats and fails to fully consider the unique spatial properties and huge data volume of holographic video. On the other hand, more research on holographic video is the opposite, usually based on the stable and low-delay characteristics of terrestrial networks, resulting in their solutions being unsuitable for the intermittent connection and significant delay characteristics of LEO channels. In addition, most existing research is application-independent, focusing on physical layer enhancement techniques such as NOMA to improve spectral efficiency, but lacking a deep understanding of the complex and long-term user QoE requirements of holographic video.

[0004] In recent years, emerging research has begun to tackle the challenges faced by this interdisciplinary frontier from different perspectives. In the field of holographic video coding, new coding standards such as Video-based Point Cloud Compression (V-PCC) and the International Standard MPEG-I for immersive video have significantly improved compression efficiency, laying the foundation for reducing data transmission volume. In terms of satellite resource management, Multi-Agent Reinforcement Learning (MARL) has been applied to inter-satellite resource collaborative allocation, while various techniques have been introduced to reduce transmission delay. Deep reinforcement learning-based algorithms have also shown great potential in energy efficiency optimization in NOMA networks. At the same time, new theoretical advances have been made in QoE modeling, including quality assessment models based on the Human Visual System (HVS) and the emerging paradigm of semantic communication. Although these research results are valuable, they have not been integrated into a unified framework to address the cross-layer, dynamic challenges faced by satellite-supported streaming transmission. However, these research advances are still fragmented: either focusing on improving a single technical layer (such as coding or resource allocation) or lacking deep adaptation to the dynamic characteristics of the satellite environment, making it difficult to integrate into a system-level solution. Most importantly, there is still a lack of a unified framework that can theoretically coordinate the design of application-layer streaming strategies with physical-layer resource allocation, thus achieving breakthroughs in two key dimensions: strict delay constraints and maximization of long-term QoE. Due to the lack of such a framework, the great potential of satellite-enabled holographic communication has not been realized. SUMMARY

[0005] Therefore, the present application aims to provide a QoE-driven low-orbit satellite holographic video cross-layer adaptive transmission method to solve the problem that the extreme data demand of holographic video transmission does not match the dynamic channel characteristics of low-orbit satellites in the prior art, resulting in transmission bottlenecks and other problems. This method addresses the transmission contradiction between the time-varying characteristics of low-orbit satellite channels and the millisecond-level latency requirements of holographic video. Through cross-layer coordination and hybrid algorithm design, stable transmission is achieved in a dynamic satellite environment. Compared with traditional methods, the long-term user experience is significantly improved.

[0006] To achieve the above object, the present application provides the following technical solutions: A QoE-driven low-orbit satellite holographic video cross-layer adaptive transmission method, specifically comprising the following steps: S1: Constructing a low-orbit satellite end-to-end holographic video streaming system model, completing holographic object space division and user field of view dynamic mapping through a 3D block model, implementing Rician fading channel modeling based on a communication model, quantifying end-to-end latency with the help of a latency model, and measuring user experience based on spatial importance and temporal stability through a QoE model; wherein QoE represents Quality of Experience; S2: design the joint decision variables of video quality level, content compression strategy and physical layer power allocation to realize the application layer and physical layer cross-layer cooperation, and convert the real-time decision sequence into a time-domain coupled continuous strategy optimization problem, which maximizes the long-term user experience quality as the target, while constraining the end-to-end delay, NOMA continuous interference cancellation feasible region and satellite power upper limit; wherein, NOMA represents non-orthogonal multiple access technology; S3: learn the joint flow strategy of quality and compression by using deep reinforcement learning, embed convex optimization to verify physical layer constraints in real time, and guarantee the feasibility of actions through the closed-loop mechanism of intelligent decision and optimization verification.

[0007] Further, in step S1, the 3D block model is that: the user moves in a three-dimensional holographic space; in each time slot , the set of blocks in the user's field of view is dynamically determined; wherein is composed of the blocks closest to the user's current position, denotes the set of all blocks of the holographic object. Considering that the delay requirement of holographic applications is extremely strict, it is assumed that the system transmits a single, high-priority block to each user in each time slot; this assumption simulates a delay-critical scenario, that is, the transmission of the most important tile for perception is prioritized, thereby allowing a focus on the core cross-layer trade-off problem in a manageable way. Such a tile is denoted as , which is selected from according to its perceptual importance; specifically, is defined as the tile in with the shortest three-dimensional Euclidean distance from the user ; the core task of the system is to determine how to select ; Each block can be encoded into different quality levels, denoted as the set ; denotes the bit rate (unit: bps) corresponding to the quality level , and satisfies ; thus, the first decision variable is , which represents the quality level of the transmission tile selected for user in time slot , satisfying:

[0008] wherein is the set of ground users, is the set of all time slots; Let ​denotes the set of all quality level selections; while the second decision is the compression strategy; let denotes a binary variable, indicating whether the block of user is compressed or not, i.e.:

[0009] where is the compression ratio; the set of all compression decisions is denoted by ; in summary, and jointly determine the data size each user needs to transmit in each time slot.

[0010] Further, in step S1, the communication model is that the channel between user and low earth orbit satellite is characterized by a Rayleigh fading model, which takes into account both strong line-of-sight (LOS) component and multipath effects; the time-varying channel coefficient is denoted by , whose gain contains small-scale fading and large-scale path loss, and dynamically changes with the movement of the satellite. It is worth noting that in the present application, users are located in the same downlink beam coverage area, but there are significant differences in geographical location. Therefore, the large-scale channel gain between users will vary significantly due to factors such as change in elevation angle, off-axis antenna gain, and quasi-static lognormal shadow fading, in addition to small-scale Rayleigh fading. This channel characteristic provides sufficient inter-user channel heterogeneity, so that power-domain non-orthogonal multiple access technology can be effectively implemented.

[0011] Under the given limited spectrum resources, the satellite adopts non-orthogonal multiple access (NOMA) technology with serial interference cancellation (SIC) to serve multiple users simultaneously; where SIC represents serial interference cancellation; in time slot , users are dynamically ordered according to their channel gains, i.e.: ; the decoding order is the inverse order of the channel strength order, i.e. from user to and up to ​In a general NOMA system, the feasibility of the SIC procedure depends on satisfying a set of cascaded signal-to-interference-plus-noise ratio (SINR) constraints. A common approach is to enhance the detectability of the signals of the users with weaker channels at the receivers of the users with stronger channels by allocating more power to them, making them easier to decode and cancel. However, as demonstrated in advanced NOMA systems, such as those in complex satellite environments, the final power allocation results tend to be the outcome of joint optimization of multiple optimization objectives, and thus do not always strictly follow this simple power allocation principle. Therefore, to maintain the simplicity and tractability of the analysis and focus on the cross-layer interaction between the high-level application layer decisions and the overall system performance, a common operational assumption can be adopted: the underlying SIC procedure is considered to be successful, and its specific cascaded SINR constraints are not explicitly enforced in the high-level optimization problem. This operational assumption enables the focus to be on the key performance bottleneck, i.e., whether each user can successfully decode its own signal after the successful execution of the SIC chain; thus, the focus can be on modeling the SINR of each user when decoding its own signal, which is given by:

[0012] where is the signal-to-interference-plus-noise ratio of user k , is the noise power spectral density, is the channel bandwidth, denotes the transmit power allocated to user in time slot , satisfying:

[0013]

[0014] where is the maximum transmit power of the satellite, is the set of ground users, is the set of all time slots; let denote the power allocation decision. To ensure that each user can reliably receive its own data under the aforementioned operational assumption, a minimum SINR threshold constraint is imposed:

[0015] where is the SINR threshold required for successful decoding; then, according to the Shannon capacity formula, the achievable data rate of user is expressed as: .

[0016] ​Furthermore, in step S1, the time delay model is: decision variable quality level Compression strategy The data load for each block is determined by the power allocation vector. This determines the achievable transmission rate. These factors collectively determine end-to-end latency, a key performance indicator for real-time holographic video streaming. (User) In the time slot The total delay of the tiles is denoted as It consists of the following four parts: (1) Compression delay If compression is selected, the quality level will be... This will generate compression latency on the satellite edge server; this latency is proportional to the original tile size, therefore:

[0017] in, It is the playback duration. It is the server's processing power (unit: bps). (2) Transmission delay : This refers to the time required to transmit (potentially compressed) blocks over a wireless link, i.e.: (3) Decompression delay On the user device side, decoding compressed tiles will cause decompression delay.

[0018] in, It is the processing power of the user equipment (unit: bps); (4) Propagation delay This refers to the time it takes for a signal to travel one way from a satellite to a user.

[0019] in, Satellite to user distance, It's the speed of light; Therefore, the total end-to-end delay is the sum of the above components:

[0020] To ensure smooth, real-time playback, the total latency must not exceed the duration of the time slot; therefore, the following constraints apply: .

[0021] Constraints will determine decision variables The coupling is the core of the feasibility of any transmission strategy.

[0022] Further, in step S1, the QoE model is: in order to evaluate the performance of the feasible transmission strategy, a user experience quality (QoE) model is constructed, which is used to quantify the user's satisfaction. The QoE index simultaneously captures the perceived quality of the transmitted video and its temporal stability. For users In time slot , its QoE is defined as:

[0023] where, is a weight factor to penalize quality fluctuation; the term represents the perceived quality of a rendered tile, which depends on its bit rate and spatial importance in the user's field of view, and is given by:

[0024] where, is a positive constant; the positive weight and are used to balance the quality contribution from spatial importance and bit rate; captures the perceived discontinuity due to quality change over time, defined as the absolute change in quality from the previous time slot: , the initial quality is assumed to be .

[0025] Further, step S2 specifically includes: the goal of the present application is to design a dynamic strategy that can make real-time decisions in terms of quality level, compression decision and power allocation to maximize the long-term system-level user experience quality (QoE). The characteristics of this problem lie in the dynamics of the low earth orbit (LEO) satellite environment and the time sequence coupling relationship between decisions in each time slot. In order to formalize this problem, first analyze its inherent complexity from the perspective of a single time slot, and then define the real sequence optimization goal.

[0026] If the system is fixed at a certain time slot , the goal is to maximize the total QoE at that moment, which can be expressed as:

[0027]

[0028]

[0029]

[0030]

[0031]

[0032]

[0033] Since P1 is a mixed integer nonlinear programming (MINLP), which is NP-hard. However, it is not enough to solve this problem by only using a myopic strategy, i.e., solving this problem per time slot. The key issue is the coupling in time: the quality in time slot is crucial for two reasons. First, it determines the quality fluctuation in the moment . As established in the QoE model in Equation , will directly penalize the user experience score. Second, becomes the reference benchmark for comparing the quality in the next moment .

[0034] Therefore, if a greedy strategy is used in time slot , only maximizing the momentary quality, it can set an unrealistic high benchmark. This can lead to severe penalties in the future, for example, when the channel conditions in the subsequent time slots inevitably deteriorate, resulting in a significant drop in quality, thus causing an unpleasant user experience. A truly intelligent strategy must have foresight, possibly sacrificing a small amount of immediate QoE to maintain stability over time and achieve higher cumulative returns in the long term. Therefore, the real goal of problem P1 is to find an optimal strategy that can map the system state to actions over time; the optimal strategy must maximize the expected cumulative QoE, which is a sequential policy optimization problem in form:

[0035] where denotes the discount factor, the expectation is taken over all possible trajectories induced by the policy , and the decision at each step must satisfy the instantaneous constraints defined in problem P1; denotes the mathematical expectation, used to calculate the long-term statistical average of system performance indicators under the influence of random factors; solving problem P2 requires a framework that can learn and plan under uncertainty, which requires the use of deep reinforcement learning (DRL).

[0036] Further, step S3 specifically comprises: using a hybrid algorithm to combine DRL with convex optimization to solve problem P2; DRL is used to handle the high-level sequential decision process, while the embedded convex optimization module checks the feasibility of low-level resource allocation at each step, where DRL denotes deep reinforcement learning; ​First, problem P2 is modeled as a Markov Decision Process (MDP) and a DRL algorithm is adopted to learn the optimal policy ; where the central agent interacts with the environment over time; its components are as follows: 1) State : the current state of the environment, denoted as: where is the three-dimensional Euclidean distance between the user and the tile ; 2) Action : the discrete decision made by the agent at the current time, denoted as: ; 3) Reward : a scalar value reflecting the quality of the action; its calculation depends on the feasibility of the action; defined as:

[0037] where is a large negative constant to penalize infeasible actions; Next, the Proximal Policy Optimization (PPO) algorithm is adopted to solve the above Markov Decision Process (MDP), where PPO represents Proximal Policy Optimization; PPO is an advanced actor-critic based deep reinforcement learning method; the PPO algorithm trains an actor network to represent the policy, while training a critic network to estimate the state value function; the actor network is updated by minimizing the clipped surrogate objective function of PPO , i.e.:

[0038] where denotes the empirical average over a finite batch of samples, is the ratio of policy probabilities, is the clipping hyperparameter to control the update step size, and is the advantage estimation discount factor and are used to reduce the variance in the policy gradient; A key component in the hybrid algorithm is to determine whether the action selected by the PPO agent is feasible; an action is considered feasible only if there exists a power allocation that satisfies all the physical layer constraints; this feasibility check relies on the constraint conditions ; for a given action The minimum data rate required for each user can be derived; first, according to the total delay constraint: Substitute each delay into the available: Thus, the minimum achievable data rate for user is defined as: where denotes the duration of a single block play. It is noted that if the denominator of the above expression is not positive, the action is obviously infeasible; otherwise, using the Shannon capacity formula, the minimum rate requirement can be directly translated into a minimum SINR target:

[0039] Therefore, the feasibility problem of the action is reduced to: whether there exists a power allocation that can simultaneously satisfy the SINR targets in the above for all users and other system constraints; this can be formalized as the following convex feasibility problem:

[0040]

[0041]

[0042]

[0043] where denotes the demand threshold representing the minimum quality of service for user k at time slot t; since all the constraints in problem P3 with respect to the power variable are linear, it constitutes a convex feasibility problem; this problem can be efficiently solved by the standard convex optimization solver CVX. The output of the solver determines whether to return a reward signal to the DRL agent. That is, the algorithm flow is as follows: First, initialize the actor network and critic network, and construct an experience replay buffer to store state-action-reward sequences; during training, for each episode, the system observes the initial state containing the low earth orbit satellite channel state, user field of view remapping information, and historical user quality of experience (QoE) and repeatedly executes the core operations in the episode time step loop ; for each time step , the joint flow policy action is output by the actor network based on the current state Secondly, the output joint action is substituted into the constraint condition of the optimization problem to verify whether the SINR and end-to-end delay meet the physical layer requirement: if the action is feasible, the QoE gain is calculated as the instant reward ; if not, the full compression mode is forced to switch first and the constraint is verified again, and if it still does not meet the requirement, the penalty reward is triggered ; ; Then, the new state of the environment feedback is observed , and the state-action-reward tuple is stored in the experience replay buffer; when the sample size of the buffer meets the batch training requirement, the sampled batch data is used to update the network and the network ; Finally, the agent continuously optimizes the policy by iteratively training the actor and critic networks until the reward converges.

[0044] The beneficial effects of the present application are as follows: the present scheme realizes the linkage optimization of holographic video spatial tile level flow strategy and satellite NOMA physical layer resource allocation through 3D block space perception and cross-layer decision design, breaks through the imbalance dilemma of traditional scheme QoE and resource efficiency; with the help of continuous strategy optimization modeling in time domain, the long-term user experience maximization is taken as the goal to break through the limitation of single time slot greedy strategy, and the industry pain point of experience fluctuation under time-varying channel is solved; relying on the hybrid algorithm of PPO deep reinforcement learning and convex optimization real-time verification, both long-term strategy learning and real-time satisfaction of physical layer constraints are utilized, which significantly improves the action feasibility compared with pure DRL scheme; finally, in the scene of low-orbit satellite time-varying channel and user field of view drift, the end-to-end delay of holographic video is stabilized at the level of milliseconds, supporting stable transmission of immersive services such as holographic conference and virtual collaboration, and filling the technical gap of collaborative transmission of high data rate holographic video and dynamic satellite channel.

[0045] Other advantages, objects, and features of the present application will be apparent to those skilled in the art from the following specification, which is to be taken in conjunction with the accompanying drawings, wherein: BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to make the objects, technical solutions and advantages of the present application clearer, the preferred embodiments of the present application will be described in detail below with reference to the accompanying drawings, in which: Figure 1 A holographic content transmission system framework model considering low-orbit satellites is provided for the present application; Figure 2 A convergence graph for joint video quality selection, compression decision and resource allocation is provided for the present application; Figure 3 A graph showing the relationship between total QoE and bandwidth for users; Figure 4 A graph showing the relationship between total QoE and slot duration for users; Figure 5 A graph showing the relationship between total QoE for users and the processing capacity of user devices. Detailed Implementation

[0047] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0048] Please see Figures 1-5 This invention provides a QoE-driven method for cross-layer adaptive transmission of low-orbit satellite holographic video, specifically including the following steps: S1: Construct an end-to-end holographic video stream transmission system model for low-Earth orbit satellites. This involves using a 3D block model to partition the holographic object space and dynamically map the user's field of view; implementing Rician fading channel modeling based on a communication model; quantifying end-to-end latency using a latency model; and integrating spatial importance and temporal stability metrics based on a QoE model to measure user experience. Specifically, this includes: The established low Earth orbit satellite communication system aims to aggregate [communication data] to ground users. It provides holographic video streaming services. The system operates in a three-dimensional space containing holographic objects located at the origin. The space is spatially divided into a set... A different 3D block, denoted as The satellite is equipped with a mobile edge computing server for caching holographic content. The system operates within a limited timeframe. Within this time range, the time interval is divided into fixed durations. Discrete time slots.

[0049] S11: 3D Block Model The user moves within a three-dimensional holographic space. (In time slots) ,user The set of blocks in the field of view It is dynamically determined. The closest to the user's current location Consider the stringent delay requirements of holographic applications, we assume that the system transmits a single, high-priority tile to each user in each time slot. This assumption models a delay-critical scenario, where the most perceptually important tile is transmitted first, thus allowing a manageable way to focus on the core cross-layer tradeoff problem. Such a tile is denoted by , which is selected from according to its perceptual importance. Specifically, we define as the tile in with the shortest 3D Euclidean distance to user . The core task of the system is to determine how to select .

[0050] Each tile can be encoded into different quality levels, denoted by the set . Let denote the bit rate (in bps) corresponding to quality level , and satisfy . The first decision variable is thus , which denotes the selected quality level of the transmitted tile for user in time slot , satisfying:

[0051] Let denote the set of all quality level selections. The second decision is the compression strategy. Let denote a binary variable, which indicates whether to compress the tile of user in time slot , i.e.,

[0052] where is the compression ratio. The set of all compression decisions is denoted by . In summary, and jointly determine the size of the data to be transmitted for each user in each time slot.

[0053] S12: Communication Model The channel between user and the LEO satellite is characterized by a Rayleigh fading model, which takes into account both the strong line-of-sight (LOS) component and multipath effects. The time-varying channel coefficient is denoted by , whose gain contains small-scale fading and large-scale path loss, and dynamically changes with the movement of the satellite. It is worth noting that in this invention, users are located in the same downlink beam coverage area, but there are significant differences in geographical location. Therefore, the large-scale channel gain between users will be significantly different due to factors such as changes in elevation angle, off-axis antenna gain, and quasi-static lognormal shadow fading, in addition to small-scale Rayleigh fading. This channel characteristic provides sufficient inter-user channel heterogeneity, enabling the effective implementation of power-domain non-orthogonal multiple access technology.

[0054] Under the given limited spectrum resources, the satellite adopts non-orthogonal multiple access (NOMA) technology with serial interference cancellation (SIC) to serve multiple users simultaneously. In the time slot , users are dynamically sorted according to their channel gains, i.e., . The decoding order is the reverse order of the channel strength order, i.e., from to to . In a general NOMA system, the feasibility of the SIC process depends on meeting a set of cascaded signal-to-interference-plus-noise ratio (SINR) constraints. A common method is to allocate more power to users with weaker channels to enhance the detectability of their signals at the receiving end of strong users, making them easier to decode and eliminate. However, as demonstrated in advanced NOMA systems, such as systems in complex satellite environments, the final power allocation result is often the result of joint optimization of multiple optimization objectives, so it does not always strictly follow this simple power allocation principle. Therefore, in order to maintain the simplicity and tractability of the analysis and focus on the cross-layer interaction between high-level application layer decisions and overall system performance, a common operating assumption can be adopted: the underlying SIC process is considered successful, and its specific cascaded SINR constraints are not explicitly enforced in the high-level optimization problem. This operating assumption can focus attention on the key performance bottleneck, i.e., whether each user can successfully decode its own signal after successfully executing the SIC chain. Therefore, we can focus on modeling the SINR of each user when decoding its own signal, which is given by:

[0055] where is the noise power spectral density, is the channel bandwidth, denotes the transmit power allocated to user in time slot , satisfying:

[0056]

[0057] where is the maximum transmit power of the satellite. Let denote the power allocation decision. To ensure that each user can reliably receive its own data under the aforementioned operational assumptions, a minimum SINR threshold constraint is imposed:

[0058] where is the SINR threshold required for successful decoding. Then, according to the Shannon capacity formula, the achievable data rate for user can be expressed as: .

[0059] S13: Latency Model Decision variable quality level and compression policy determines the data load of each block, while the power allocation vector determines the achievable transmission rate . These factors together determine the end-to-end delay, which is a key performance indicator for real-time holographic video streaming. The total delay of a tile for user in time slot is denoted by and consists of the following four parts: (1) Compression delay : If compression is chosen (i.e., ), a compression delay is incurred at the satellite edge server. This delay is proportional to the original tile size, so we have:

[0060] where is the playback duration, is the server’s processing capability (in bps).

[0061] (2) Transmission delay : This refers to the time required to transmit the (possibly compressed) block over the wireless link, i.e.,

[0062] (3) Decompression delay : At the user device, decoding the compressed tile incurs a decompression delay:

[0063] where is the user device’s processing capability (in bps).

[0064] (4) Propagation delay the time it takes for a signal to travel from the satellite to the user one way, i.e.,

[0065] where is the distance from the satellite to the user, is the speed of light.

[0066] Thus, the total end-to-end delay is the sum of the above parts:

[0067] To ensure smooth, real-time playback, the total delay must not exceed the slot duration. Therefore, there is the following constraint:

[0068] The constraints couple the decision variables together, and are the core of the feasibility of any transmission strategy.

[0069] S14: QoE model To evaluate the performance of a feasible transmission strategy, a quality of experience (QoE) model is constructed to quantify the user’s satisfaction. The QoE index captures both the perceived quality of the transmitted video and its temporal stability. For a user at slot , its QoE is defined as:

[0070] where is a weight factor to penalize quality fluctuations. The term represents the perceived quality of a rendered tile, which depends on its bitrate and spatial importance in the user’s field of view, and is given by:

[0071] where is a positive constant, and the positive weights and are used to balance the quality contributions from spatial importance and bitrate. captures the perceived discontinuity due to quality changes over time, and is defined as the absolute change in quality from the previous slot: with the initial quality assumed to be .

[0072] ​S2: design the joint decision variables of video quality level, content compression strategy, and physical layer power allocation to realize the cross-layer coordination between application layer and physical layer, and transform the real-time decision sequence into a time-domain coupled continuous strategy optimization problem, which aims to maximize the long-term user experience while constraining the end-to-end delay, NOMA successive interference cancellation feasible region and satellite power upper limit. Specifically, it includes: The object of the present application is to design a dynamic strategy that can make real-time decisions in terms of quality level, compression decision and power allocation to maximize the long-term system-level user experience quality (QoE). The characteristics of this problem lie in the dynamics of the low earth orbit (LEO) satellite environment and the time sequence coupling relationship between decisions in each time slot. In order to formalize this problem, first analyze its inherent complexity from the perspective of a single time slot, and then define the real sequence optimization target.

[0073] If the system is fixed at a certain time slot , the target will be to maximize the total QoE at that moment, which can be expressed as:

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080] Since this problem is a mixed integer nonlinear programming (MINLP), it is an NP-hard problem. However, it is not enough to use only short-sighted strategy, i.e. solving the problem per time slot. The key problem lies in the coupling in time: the perceived quality in time slot is crucial for the following two reasons. First, it determines the immediate quality fluctuation . As established in the QoE model in formula , will directly penalize the user experience score. Second, becomes the reference benchmark for comparing the quality of the next moment.

[0081] Therefore, if the time slot Adopting a greedy strategy that maximizes only the immediate quality can set an unrealistic high bar. This can lead to severe penalties in the future, for example, when the channel conditions of the subsequent time slots inevitably deteriorate, resulting in a significant drop in quality and thus an unpleasant user experience. A truly intelligent strategy must be farsighted, possibly sacrificing a small amount of immediate QoE to maintain stability over time and achieve higher cumulative benefits in the long term. Therefore, the real goal is to find an optimal strategy that can map the system state to actions over time. The optimal strategy must maximize the expected cumulative QoE, which is formally a sequential policy optimization problem:

[0082] where is the discount factor, and the expectation is taken over all possible trajectories induced by the policy . The decision at each step must satisfy the instantaneous constraints defined in problem P1. Solving problem P2 requires a framework that can learn and plan under uncertainty, which requires the use of deep reinforcement learning (DRL).

[0083] S3: Learn the joint flow strategy of quality and compression using DRL, embed convex optimization to verify physical layer constraints in real time, and guarantee the feasibility of actions through a closed-loop mechanism of intelligent decision and optimization verification. Specifically, it includes: A hybrid algorithm is used to combine deep reinforcement learning with convex optimization to solve problem P2. DRL is used to handle the high-level sequential decision-making process, while the embedded convex optimization module checks the feasibility of low-level resource allocation at each step.

[0084] First, the problem can be modeled as a Markov decision process (MDP), and a DRL algorithm is used to learn the optimal policy . The central agent interacts with the environment over time. Its components are as follows: 1) State : The current state of the environment, represented as: where is the three-dimensional Euclidean distance between the user and the tile ; 2) Action : The discrete decision made by the agent at the current time, represented as: ; 3) Reward : A scalar value that reflects the quality of the action. Its calculation depends on the feasibility of the action. defined as:

[0085] where is a large negative constant to penalize infeasible actions.

[0086] Next, the proximal policy optimization (PPO) algorithm is employed to solve the above Markov decision process (MDP). PPO is an advanced actor-critic based deep reinforcement learning method. The algorithm trains an actor network to represent the policy, while training a critic network to estimate the state value function. The actor network is updated by minimizing the clipped surrogate objective function of PPO , i.e.,

[0087] where denotes the empirical average over a finite batch of samples, is the ratio of the policy probabilities, is a clipping hyperparameter to control the update step size, and is the advantage estimation discount factor and are used to reduce the variance in the policy gradient.

[0088] A key component in the hybrid algorithm is to determine whether the action selected by the PPO agent is feasible. An action is considered feasible only if there exists a power allocation that satisfies all the physical layer constraints. This feasibility check relies on the constraint . For a given action , the minimum data rate required by each user can be derived. First, according to the total delay constraint: substitute the individual delays into the achievable: from which the minimum achievable data rate for user is obtained, defined as: It is noted that if the denominator of the above expression is not positive, the action is obviously infeasible. Otherwise, using the Shannon capacity formula, the minimum rate requirement can be directly translated into the minimum SINR target:

[0089] Therefore, the feasibility problem of action is reduced to: whether there exists a power allocation , which can simultaneously satisfy the SINR targets of all users in the above and other system constraints. This can be formalized as the following convex feasibility problem:

[0090]

[0091]

[0092]

[0093] Since all the constraints in problem P3 with respect to the power variables are linear, it constitutes a convex feasibility problem. This problem can be solved efficiently by the standard convex optimization solver CVX. The output of the solver determines whether to return a reward signal to the DRL agent. That is, the algorithm flow is as follows: First, the actor network and critic network are initialized, and an experience replay buffer is constructed to store state-action-reward sequences; during training, for each episode, the system observes the initial state containing the low-orbit satellite channel state, user field-of-view shift information, and historical user quality of experience (QoE) and repeatedly performs the core operations in the episode time step loop . For each time step , the actor network outputs a joint flow policy action based on the current state ; Second, the output joint action is substituted into the constraint conditions of the optimization problem to verify whether the SINR and end-to-end delay meet the physical layer requirements: if the action is feasible, the QoE gain is calculated as the immediate reward ; if it is not feasible, it is forced to switch to the full compression mode and the constraint is verified again, and if it still does not meet the requirements, a penalty reward is triggered ; Then, the new state observed by the environment feedback is stored in the experience replay buffer; when the sample size of the buffer meets the batch training requirements, the network and the network are sampled. Finally, the agent iteratively trains the actor and critic networks to continuously optimize the policy until the reward converges.

[0094] Figure 2 The convergence of the algorithm proposed in the present application is demonstrated. From Figure 2It can be observed that the algorithm experiences a rapid performance improvement in the early stage of training, and gradually converges and reaches a stable state as the training process progresses. The fluctuations during training gradually decrease, indicating that the algorithm has good convergence and stability. The entire convergence process verifies the effectiveness of the PPO algorithm in solving the optimization problem described in this paper.

[0095] Figure 3 , Figure 4 , Figure 5 The performance of the joint video quality selection, compression decision, and resource allocation scheme proposed in this invention is compared with three baselines. It can be seen that due to the flexible joint design of video quality selection, compression decision, and resource allocation, the algorithm proposed in this invention is superior to the baseline. Among them, the baseline scheme is, baseline 1 uses the PPO scheme without time smoothing mechanism, which is used to verify the ability of the agent to optimize the long-term user experience quality (QoE). The structure is exactly the same as the algorithm proposed in this invention, but a modified reward function is used during training, only considering the immediate perceptual quality. Specifically, the reward signal of the agent is , completely ignoring the time fluctuation penalty term . Baseline 2 is a PPO scheme without compression mechanism, which is used to evaluate the necessity of the compression mechanism in the resource-limited scene. This scheme uses a PPO agent to learn to select the quality level , but is limited to always transmitting without compression. That is, the compression decision is fixed to , for all users and all time slots . The agent is trained in an action space of size , and the training process is the same as the algorithm in this invention. Baseline 3 is a regularized compression + PPO quality decision scheme, which uses a separate design: the compression decision is determined by a fixed rule, and the quality level is learned by a PPO agent. Specifically, the system always enables compression for all users (i.e. ), and only uses a PPO agent to optimize the quality level selection .

[0096] Specifically, Figure 3The overall QoE of all schemes is shown to improve with increasing bandwidth. This is because larger bandwidths provide higher transmission capacity, enabling the system to transmit higher quality tiles while satisfying the delay constraint. Our proposed scheme consistently outperforms all three baseline schemes across the entire bandwidth range. Baseline 2 (no compression scheme) improves at a slower rate and remains at a lower level across all bandwidth values, as it is limited to always transmitting uncompressed tiles, which requires more bandwidth and frequently violates the timing constraint. Baseline 3 (rule-based compression scheme) performs the worst at low bandwidths, as its fixed threshold-based compression strategy fails to effectively adapt to severely constrained scenarios. Baseline 1 (no timing smoothing scheme) outperforms the other two baselines but still falls short of our proposed method, as it only optimizes instantaneous quality without considering timing stability.

[0097] Figure 4 The impact of slot duration on QoE is shown. All schemes exhibit performance improvement when the slot is longer, as this relaxes the strict delay constraint. Baselines 2 and 3 perform poorly when the slot duration is extremely short. Baseline 2 lacks the flexibility to reduce transmission time through compression, while baseline 3's rule-based compression strategy is too rigid for tight timing constraints. Baseline 1 exhibits moderate performance but its growth curve is not as smooth compared to our proposed method. Our scheme demonstrates consistent and steady improvement across the entire range, as the DRL agent learns to balance between immediate quality and timing smoothness. When the slot duration is short, many actions violate the end-to-end delay constraint and are rejected by the feasibility check, which explains the sharp drop in QoE for baseline schemes.

[0098] Figure 5 The impact of user device processing capability on QoE is shown. It can be seen that the proposed scheme exhibits monotonic improvement with increasing device processing capability, as faster decompression speeds reduce overall delay, enabling the selection of higher quality video streams. Baseline 2 (no compression scheme) exhibits a flat performance curve, as without compression, decompression delay is not generated regardless of user device processing capability, making this parameter completely irrelevant to its strategy. Baseline 3 (rule-based compression scheme) also exhibits relatively stable but flat performance, as its rule-based compression strategy is decoupled from the processing constraints at the device end. It is worth noting that baseline 1 (no timing smoothing scheme) exhibits significant oscillation, with its performance fluctuating dramatically across different user device processing capabilities. This variability stems from the short-sighted agent: it takes advantage of immediate delay relief to jump to higher quality, which in turn triggers a large quality fluctuation penalty when subsequent environmental conditions inevitably worsen.

[0099] Figure 2The convergence of the algorithm proposed in the present application is demonstrated. From Figure 2 it can be observed that the algorithm experiences rapid performance improvement in the early stage of training, and gradually converges and reaches a stable state as the training process progresses. The fluctuations during training gradually decrease, indicating that the algorithm has good convergence and stability. The entire convergence process verifies the effectiveness of the PPO algorithm in solving the optimization problem described in this paper.

[0100] Figure 3 , Figure 4 , Figure 5 The performance of the joint video quality selection, compression decision, and resource allocation scheme proposed in the present application is compared with three baselines. It can be seen that due to the flexible joint design of video quality selection, compression decision, and resource allocation, the algorithm proposed in the present application is superior to the baselines. Among them, baseline 1 uses a PPO scheme without a time smoothing mechanism, which is used to verify the ability of the agent to optimize long-term user experience quality (QoE). Its structure is exactly the same as the algorithm proposed in the present application, but a modified reward function is used during training, only considering the immediate perceptual quality. Specifically, the reward signal of the agent is , completely ignoring the time fluctuation penalty term . Baseline 2 is a PPO scheme without compression mechanism, which is used to evaluate the necessity of compression mechanism in resource limited scenarios. This scheme uses a PPO agent to learn to select quality levels , but is limited to always transmitting without compression. That is, the compression decision is fixed as , for all users and all time slots . The agent is trained in an action space of size , and the training process is the same as that of the algorithm of the present application. Baseline 3 is a regularized compression + PPO quality decision scheme, which uses a separate design: the compression decision is determined by a fixed rule, and the quality level is learned by a PPO agent. Specifically, the system always enables compression for all users (i.e. ), and only uses a PPO agent to optimize quality level selection .

[0101] Specifically, Figure 3The total QoE of all schemes is shown to improve with increasing bandwidth. This is because larger bandwidths provide higher transmission capacity, enabling the system to transmit higher quality tiles while satisfying the delay constraint. The proposed scheme outperforms all three baseline schemes consistently across the entire bandwidth range. Baseline 2 (no compression scheme) improves at a slower pace and remains at a lower level across all bandwidth values, as it is limited to always transmitting uncompressed tiles, which requires more bandwidth and frequently violates the timing constraint. Baseline 3 (rule-based compression scheme) performs worst at low bandwidths, as its fixed threshold-based compression strategy fails to adapt effectively to severely constrained scenarios. Baseline 1 (no timing smoothing scheme) outperforms the other two baselines but still falls short of our proposed method, as it only optimizes instantaneous quality without considering timing stability.

[0102] Figure 4 The impact of slot duration on QoE is shown. All schemes exhibit performance improvement when the slot is longer, as this relaxes the strict delay constraint. Baselines 2 and 3 perform poorly when the slot duration is extremely short. Baseline 2 lacks the flexibility to reduce transmission time through compression, while baseline 3's rule-based compression strategy is too rigid for tight timing constraints. Baseline 1 exhibits moderate performance but its growth curve is not as smooth compared to our proposed method. Our scheme demonstrates consistent and steady improvement across the entire range, as the DRL agent learns to balance between immediate quality and timing smoothness. When the slot duration is short, many actions violate the end-to-end delay constraint and are rejected by the feasibility check, which explains the sharp drop in QoE for baseline schemes.

[0103] Figure 5 The impact of user device processing capability on QoE is shown. It can be seen that the proposed scheme exhibits monotonic improvement with increasing device processing capability, as faster decompression speeds reduce overall delay, enabling the selection of higher quality video streams. Baseline 2 (no compression scheme) exhibits a flat performance curve, as without compression, decompression delay is independent of user device processing capability, making this parameter completely irrelevant to its strategy. Baseline 3 (rule-based compression scheme) also exhibits relatively stable but flat performance, as its rule-based compression strategy is decoupled from the processing constraints at the device end. It is worth noting that baseline 1 (no timing smoothing scheme) exhibits significant oscillation, with its performance fluctuating dramatically across different user device processing capabilities. This variability stems from the short-sighted agent: it takes advantage of immediate delay relief to jump to higher quality, which in turn triggers a large quality fluctuation penalty when subsequent environmental conditions inevitably deteriorate.

[0104] Finally, it is to be explained that the above embodiments are only used to illustrate the technical solutions of the present application but not to limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions, and all should be covered in the scope of the claims of the present application.

Claims

1. A QoE-driven method for cross-layer adaptive transmission of low-orbit satellite holographic video, characterized in that, The method includes the following steps: S1: Construct an end-to-end cross-layer holographic video stream transmission system model for low-orbit satellites. Complete the spatial division of holographic objects and dynamic mapping of user field of view through a 3D block model. Implement Rician fading channel modeling based on the communication model. Quantize end-to-end latency using a latency model. Based on the QoE model, integrate spatial importance and temporal stability to measure user experience; where QoE represents user experience quality. S2: Design joint decision variables for video quality level, content compression strategy and physical layer power allocation to achieve cross-layer collaboration between application layer and physical layer, and transform the real-time decision sequence into a temporally coupled continuous strategy optimization problem. The problem aims to maximize long-term user experience quality, while constraining end-to-end latency, the feasible region of NOMA continuous interference cancellation and satellite power limit; where NOMA represents non-orthogonal multiple access technology. S3: Utilizes a joint flow strategy of learning quality and compression through deep reinforcement learning, embeds convex optimization to verify physical layer constraints in real time, and ensures the feasibility of actions through a closed-loop mechanism of intelligent decision-making and optimization verification.

2. The QoE-driven cross-layer adaptive transmission method for low-orbit satellite holographic video according to claim 1, characterized in that, In step S1, the 3D segmented model is: the user moves in a three-dimensional holographic space; in time slots ,user The set of blocks in the field of view It is dynamically determined; among which The closest to the user's current location Composed of blocks, Represents the set of all blocks of a holographic object; Defined as In and with users The block with the shortest three-dimensional Euclidean distance; The core task of a low-Earth orbit satellite end-to-end holographic video streaming system is to determine how to transmit the video. ; Each block is encoded as A set of three different quality levels. ; Indicates quality level The corresponding bit rate, and satisfying ; Thus, the first decision variable is , indicating in time slot For users The selected transport block quality level satisfies: in, For ground users, For the set of all time slots; make Let represent the set of all quality level choices; and the second decision is the compression strategy; let Represent a binary variable used to indicate the time slot For users The blocks are compressed, that is: in, It is the compression ratio; the set of all compression decisions is denoted as . In summary, and Together they determine the amount of data each user needs to transmit in each time slot.

3. The QoE-driven low-orbit satellite holographic video cross-layer adaptive transmission method according to claim 2, characterized in that, In step S1, the communication model is: User The channel with low Earth orbit satellites is characterized by a Rayleigh fading model that considers both strong line-of-sight components and multipath effects; the time-varying channel coefficients are denoted as... Its gain It includes small-scale fading and large-scale path loss, and changes dynamically as the satellite moves; Given limited spectrum resources, the satellite employs NOMA technology with SIC (Sequential Interference Cancellation) to simultaneously serve multiple users; where SIC represents Sequential Interference Cancellation; in time slots Users are dynamically sorted according to their channel gain, i.e.: Its decoding order is the reverse of the channel strength order, that is, from the user... arrive Until Assuming we can focus on the key performance bottleneck—whether each user can successfully decode its own signal after receiving signals from other users and successfully executing the SIC chain—we can then focus on each user's performance. The SINR model when decoding its own signal is given by the following formula: in, For users k Signal-to-interference-to-noise ratio, For noise power spectral density, For channel bandwidth, Indicates in time slot Assigned to user The transmission power satisfies: in, That is the satellite's maximum transmission power. For ground users, Let be the set of all time slots; This represents the power allocation decision; assuming each user can reliably receive its own data, a minimum SINR threshold constraint is imposed: in, It is the SINR threshold required for successful decoding; then, according to the Shannon capacity formula, the user... The achievable data rate is expressed as: .

4. The QoE-driven cross-layer adaptive transmission method for low-orbit satellite holographic video according to claim 3, characterized in that, In step S1, the latency model is: User In the time slot The total delay of the tiles is denoted as It consists of the following four parts: (1) Compression delay If compression is selected, the quality level will be... This will generate compression latency on the satellite edge server; this latency is proportional to the original tile size, therefore: in, It is the playback duration. It refers to the server's processing power; (2) Transmission delay : refers to the time required for a block to be transmitted over a wireless link, i.e.: (3) Decompression delay On the user device side, decoding compressed tiles will cause decompression delay. in, It refers to the processing power of user equipment; (4) Propagation delay This refers to the time it takes for a signal to travel one way from a satellite to a user. in, Satellite to user distance, It's the speed of light; Therefore, the total end-to-end delay is the sum of the above components: To ensure smooth, real-time playback, the total latency must not exceed the duration of the time slot; therefore, the following constraints apply: 。 5. The QoE-driven low-orbit satellite holographic video cross-layer adaptive transmission method according to claim 4, characterized in that, In step S1, the QoE model is: for users In the time slot Its QoE is defined as: in, It is a weighting factor used to penalize quality fluctuations; Item The perceptual quality of a rendered tile depends on its bit rate and spatial importance within the user's field of view, and is given by the following formula: in, It is a positive constant; positive weight. and Used to balance the quality contribution of space importance and bit rate; The perceived discontinuity caused by changes in mass over time was captured and defined as the absolute change in mass in the previous time slot: The initial mass assumption is: .

6. The QoE-driven low-orbit satellite holographic video adaptive transmission method according to claim 5, characterized in that, Step S2 specifically includes: if the system is in a specific time slot Once fixed, the objective becomes maximizing the instantaneous total QoE, expressed as: The real goal of problem P1 is to find an optimal strategy. This strategy maps system states to time-varying actions; the optimal strategy must maximize the expected cumulative QoE, which is formally a serialized policy optimization problem. in, This represents the discount factor, and the expected value is in the strategy. The decision is made on all possible trajectories induced, and each step must satisfy the instantaneous constraints defined in problem P1. It represents the mathematical expectation and is used to calculate the long-term statistical average of system performance indicators under the influence of random factors.

7. The QoE-driven low-orbit satellite holographic video cross-layer adaptive transmission method according to claim 6, characterized in that, Step S3 specifically includes: using a hybrid algorithm that combines DRL with convex optimization to solve problem P2; DRL is used to handle the high-level sequential decision-making process, while the embedded convex optimization module checks the feasibility of low-level resource allocation at each step, where DRL stands for Deep Reinforcement Learning. First, problem P2 is modeled as a Markov decision process, and the DRL algorithm is used to learn the optimal policy. The central agent interacts with the environment over time; its components are as follows: 1) Status The current state of the environment is represented as: ,in, User With tiles The three-dimensional Euclidean distance between them; 2) Actions The discrete decision made by the agent at the current moment is represented as: ; 3) Rewards A scalar value that reflects the quality of an action; its calculation depends on the feasibility of the action. Defined as: in, It is a negative constant used to penalize infeasible actions; Next, the PPO algorithm is used to solve the above Markov decision process, where PPO stands for proximal policy optimization; PPO is an advanced actor-critic-based deep reinforcement learning method; the PPO algorithm is used to train an actor network. To represent the strategy, while training a network of critics. Used to estimate the state value function; the actor network minimizes the truncated agent objective function of PPO. To update, that is: in, This represents the empirical average over a finite number of sample batches. It is the strategy probability ratio. It is a truncation hyperparameter used to control the update step size. It is the advantage estimation discount factor and Used to reduce variance in policy gradients; Determine the action selected by the PPO agent. Is it feasible; an action Only when there exists a power allocation that satisfies all physical layer constraints It is only considered feasible at a certain time; this feasibility check depends on the constraints. For a given action The minimum data rate required for each user is derived; firstly, based on the total latency constraint: Substituting the various delays, we get: From this, we can obtain the user's information. Minimum achievable data rate Defined as: ,in, This represents the playback duration of a single block; if the denominator of the above expression is not positive, then the action... This is clearly not feasible; otherwise, using Shannon's capacity formula, the minimum rate requirement can be set... Directly convert to the minimum SINR target: Therefore, action The feasibility problem simplifies to: Does a power allocation exist? This can simultaneously satisfy all users' SINR targets and other system constraints mentioned above; this can be formalized as the following convex feasibility problem: in, This represents the minimum service quality demand threshold for user k in time slot t, due to all the power variables in problem P3. The constraints are all linear, thus constituting a convex feasibility problem; this problem is solved efficiently by the standard convex optimization solver CVX.