Video Streaming Dual-Time-Scale Wireless Transmission Method, System, and Storage Medium

Through dual-time scale scheduling and cross-layer optimization strategies, combined with deep reinforcement learning networks, dynamic allocation of bandwidth and power is solved, and the problem of channel quality changes in streaming video transmission is improved, and user experience and resource utilization are improved.

CN119946329BActive Publication Date: 2025-07-22UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510422219.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing streaming video transmission and wireless communication systems are difficult to flexibly respond to changes in dynamic channel quality, which makes it difficult to ensure the quality of user experience, and the resource utilization rate is not high, making it impossible to take into account low latency, high throughput and fairness.

Method used

Using dual-time scale scheduling and cross-layer optimization strategies, video frames are divided into basic layer and enhancement layer through scalable video encoding, combining the scheduling and utility agents of deep reinforcement learning network architecture and resolution scaling agents to dynamically allocate bandwidth and power, and realize cross-layer state observation and double-level reward feedback.

Benefits of technology

Reduce the impact of channel fluctuations on a small time scale, adaptively adjust video resolution on a large time scale, improve user experience quality and system resource utilization efficiency, and take into account the worst user frame success rate and average resolution and time delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946329B_ABST
    Figure CN119946329B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of communication technologies, and discloses a video streaming dual-time-scale wireless transmission method, system, and storage medium. The method includes: dividing each video frame in a video stream into a base layer and multiple enhancement layers based on scalable video coding technology; dynamically allocating bandwidth and power for a user through a scheduling and utility agent on a time slot-level time scale, and dynamically selecting the video frame resolution to be transmitted to the user through a resolution scaling agent on a video frame-level time scale; the scheduling and utility agent and the resolution scaling agent are jointly trained using a separate deep reinforcement learning network architecture, and a cross-layer state observation mechanism and a two-level reward feedback mechanism for the application layer and the transport layer are proposed; through dual-time-scale scheduling combined with deep reinforcement learning, the present invention effectively reduces the influence of channel fluctuations on a small time scale and adaptively adjusts the video resolution on a large time scale.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and particularly to a dual-time-scale wireless transmission method, system, and storage medium for video streaming media. Background Art

[0002] Existing streaming video transmission and wireless communication systems are usually designed independently. The video end can only distribute content based on the fixed pipelines provided by operators, and it is difficult to flexibly cope with dynamic channel quality changes, resulting in the difficulty of fully guaranteeing the quality of experience (QoE) of users.

[0003] On the other hand, due to the lack of close coordination between the underlying wireless resource scheduling and the upper-layer video encoding and decoding strategies, the resource utilization rate is often not high, and it is impossible to fully meet the concurrent needs of multiple users. Most existing studies optimize local links, such as achieving adaptive adjustment for a single encoding method or transport layer protocol, but ignore the end-to-end joint design of the complete system, and it is impossible to balance fairness and network efficiency while ensuring low latency and high throughput.

[0004] Nowadays, with the gradual introduction of virtualization and software-defined technologies in 5G and future network architectures, the programmability and scalability of communication systems have been significantly improved, and the combination of edge computing and intelligent control has made cross-layer joint optimization possible. Based on this, comprehensive optimization of cross-layer perception and cooperative scheduling at the overall system level has become the key breakthrough for further improving video quality and resource utilization rate. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a dual-time-scale wireless transmission method, system, and storage medium for video streaming media. In the case of large and small time-scale fluctuations in the wireless transmission channel, aiming at the problem that existing solutions are difficult to balance user experience and network resource efficiency, a combination of dual-time-scale scheduling and cross-layer optimization strategies is proposed to achieve high-quality video transmission and system performance improvement.

[0006] To solve the above technical problems, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a dual-time-scale wireless transmission method for video streaming media, including:

[0008] Based on scalable video coding technology, each video frame in the video stream is divided into a base layer and multiple enhancement layers, and the corresponding relationship between the video frame resolution level transmitted to the user and the number of enhancement layers is established;

[0009] Construct a dual-time-scale decision-making mechanism: Dynamically allocate bandwidth and power for users through a scheduling and utility agent at the time-slot level time scale, and dynamically select the video frame resolution for users through a resolution scaling agent at the video-frame level time scale;

[0010] The scheduling and utility agent and the resolution scaling agent are jointly trained using a separate deep reinforcement learning network architecture;

[0011] A cross-layer state observation mechanism for the application layer and the transport layer: Construct an application layer state with a video-frame period as the observation granularity and a transport layer state with a time-slot as the observation granularity; When the transport layer state is input to the resolution scaling agent, it needs to be converted into the average value of multiple time-slot states within the target video-frame period;

[0012] A two-level reward feedback mechanism: The reward function of the scheduling and utility agent includes an instantaneous resource efficiency metric and a cross-frame transmission progress penalty term, where the resources include the bandwidth and power; The reward function of the resolution scaling agent is based on the video-frame transmission result, and calculates the weighted value of the video-frame success rate, average resolution level, delay penalty, and resolution degradation penalty of the worst user.

[0013] In one embodiment, the scheduling and utility agent includes a bandwidth decision network, a power decision network, a resource evaluation network, and a target resource evaluation network; The bandwidth decision network and the power decision network are collaboratively optimized under the guidance of the same resource evaluation network;

[0014] The resolution scaling agent includes a resolution decision network, a resolution evaluation network, and a target resolution evaluation network; The resolution decision network is optimized under the guidance of the resolution evaluation network.

[0015] In one embodiment, the construction of the application layer state with a video-frame period as the observation granularity and the transport layer state with a time-slot as the observation granularity specifically includes:

[0016] The application layer state includes a video-frame delay sequence, a resolution level history, a video-frame success indicator, and a resolution decision action history; The transport layer state includes a queue length, a network throughput, a channel attenuation sequence, and a resource allocation action history.

[0017] In one embodiment, the cross-layer state observation mechanism for the application layer and the transport layer further includes:

[0018] Before the application layer state and the transport layer state are input into the scheduling and utility agent and the resolution scalability agent, the gated recurrent unit is used to process the most recent consecutive set of historical values of the application layer state and the transport layer state to extract the corresponding temporal features; among them, the transport layer state uses an observation sequence with a time slot as the granularity when the scheduling and utility agent makes a decision, and is averaged within the video frame period when the resolution scalability agent makes a decision to form a state average with the video frame as the granularity; therefore, the video frame in the resolution scalability agent The mean value of the transport layer state and the time slot in the scheduling and utility agent The transport layer state of Satisfy the relationship:

[0019] ;

[0020] Among them, Indicates the set of time slot indices within the video frame , Indicates the number of time slots included in this video frame period, Is the time slot index; the application layer state uses the state value with the same video frame time granularity when the scheduling and utility agent makes a decision and when the resolution scalability agent makes a decision.

[0021] In one of the embodiments, the two-level reward feedback mechanism specifically includes:

[0022] The reward in the scheduling and utility agent Is:

[0023] ;

[0024] Among them, Are the coefficients of the spectral efficiency reward, the power efficiency reward, the maximum remaining required rate user penalty, and the video frame transmission success reward item or the video frame transmission failure penalty item respectively; And Are the bandwidth ratio and the power ratio allocated to the user In the time slot respectively, And Are the spectral efficiency and the power efficiency respectively, And Are the remaining data volume and the remaining time of the user In the current video frame, Is the video frame transmission success indicator that takes effect at the end of each video frame, Is the user set;

[0025] The reward of the resolution scalability agent Is:

[0026] ;

[0027] Among them, are the weight parameters of the video frame transmission success reward, the video frame resolution reward, the video frame resolution downgrade penalty, and the video frame delay penalty, respectively; is the number of video frames that the user should transmit in one round of training; and are the number of video frame downgrade layers and the video frame delay, respectively, is the number of users in the user set ; is the total number of video frames in the video stream, represents the video frame of user resolution level.

[0028] In one embodiment, the dynamic allocation of bandwidth and power for users by scheduling and utility agents on a time-slot level time scale, and the dynamic selection of the video frame resolution transmitted for users by a resolution scaling agent on a video frame level time scale specifically include:

[0029] Before the start of each time slot, the scheduling and utility agent receives the application layer state and the transport layer state processed by time series, and the output actions include, on the premise of satisfying the maximum bandwidth constraint and the maximum power constraint, the bandwidth ratio allocated to user and the power ratio in time slot :

[0030] ;

[0031] ;

[0032] The bandwidth and power allocated to user are respectively:

[0033] ,, ;

[0034] is the maximum available bandwidth for allocation, represents the maximum available power for allocation;

[0035] At the end of each video frame transmission, the resolution scaling agent receives the application layer state and the transport layer state processed by time series, and the output action is the encoded value of the N-bit L-value digital code of all user resolution decision actions:

[0036] ;

[0037] wherein, represents the video frame selected by the user from resolution levels, being the resolution of the video frame, means not transmitting the video frame, means transmitting the base layer, means the base layer and enhanced layers.

[0038] In one embodiment, the scheduling and utility agent and the resolution scalability agent are jointly trained using a separate deep reinforcement learning network architecture, specifically including:

[0039] Before each time slot, the scheduling and utility agent, based on the state passes through the bandwidth decision network and the power decision network of the scheduling and utility agent to obtain the action . After implementing the action, the reward of the current time slot and the new state after implementing the action are obtained. is stored in the buffer of the scheduling and utility agent. After batch_size time steps, network optimization updates are performed: the state is input into the target resource evaluation network. Based on the output and the reward, the advantage function of the scheduling and utility agent is calculated. Based on the advantage function and the reward, the bandwidth decision network and the power decision network within the scheduling and utility agent are updated; based on the advantage function and the output of the resource evaluation network, the evaluation loss is calculated, and the resource evaluation network is updated based on the evaluation loss; the target resource evaluation network is periodically updated based on the resource evaluation network;

[0040] Before the transmission video frame period, the resolution scalability agent, based on the state passes through the resolution decision network of the resolution scalability agent to obtain the action . After implementing the action, the reward of the video frame and the new state after implementing the action are obtained. is stored in the buffer of the resolution scalability agent. After batch_size time steps, network optimization updates are performed: the state is input into the target resolution evaluation network. Based on the output and the reward, the advantage function of the resolution scalability agent is calculated. Based on the advantage function and the reward, the resolution decision network within the resolution scalability agent is updated; based on the advantage function and the output of the resolution evaluation network, the evaluation loss is calculated, and the resolution evaluation network is updated based on the evaluation loss; the target resolution evaluation network is periodically updated based on the resolution evaluation network.

[0041] In a second aspect, the present invention provides a video streaming dual-time-scale wireless transmission system, including:

[0042] A video frame layering module: Based on scalable video coding technology, each video frame in the video stream is divided into a base layer and multiple enhancement layers, and the corresponding relationship between the video frame resolution level transmitted to the user and the number of enhancement layers is established;

[0043] A dual-time-scale decision module: Dynamically allocates bandwidth and power for users through scheduling and utility agents on the time slot-level time scale, and dynamically selects the video frame resolution transmitted to the user through a resolution scaling agent on the video frame-level time scale;

[0044] A training module: The scheduling and utility agents and the resolution scaling agent are jointly trained using a separate deep reinforcement learning network architecture;

[0045] A cross-layer state observation module: Constructs an application layer state with a video frame period as the observation granularity and a transport layer state with a time slot as the observation granularity; when the transport layer state is input to the resolution scaling agent, it needs to be converted into the average value of multiple time slot states within the target video frame period;

[0046] A two-level reward feedback module: The reward function of the scheduling and utility agent includes an instantaneous resource efficiency metric and a cross-frame transmission progress penalty term, where the resources include the bandwidth and power; the reward function of the resolution scaling agent is based on the video frame transmission result, and calculates the weighted value of the video frame success rate, average resolution level, delay penalty, and resolution degradation penalty of the worst user.

[0047] In one embodiment, the scheduling and utility agent includes a bandwidth decision network, a power decision network, a resource evaluation network, and a target resource evaluation network; the bandwidth decision network and the power decision network are collaboratively optimized under the guidance of the same evaluation network;

[0048] The resolution scaling agent includes a resolution decision network, a resolution evaluation network, and a target resolution evaluation network; the resolution decision network is optimized under the guidance of the resolution evaluation network.

[0049] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in any one of the embodiments in the first aspect are implemented.

[0050] Compared with the prior art, the beneficial technical effects of the present invention are:

[0051] Through dual-time-scale scheduling combined with deep reinforcement learning, the present invention effectively reduces the impact of channel fluctuations on a small time scale and adaptively adjusts the video resolution on a large time scale, thus taking into account key metrics such as the worst-user frame success rate, average resolution, fairness among users, and latency, and significantly improving the quality of user experience and the utilization efficiency of system resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a flowchart of the method in an embodiment of the present invention.

[0053] Figure 2 It is a framework schematic diagram of the dual-time-scale wireless transmission method used in an embodiment of the present invention.

[0054] Figure 3 It is a schematic diagram of the control effect to be achieved in an embodiment of the present invention.

[0055] Figure 4 It is a schematic diagram of the intelligent agent training architecture and process to be used in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] A preferred embodiment of the present invention will be described in detail below with reference to the accompanying drawings.

[0057] The present invention optimizes wireless resources on a small time scale, "cuts peaks and fills valleys" for short-term channel fluctuations, and provides relatively stable throughput capacity; and adaptively adjusts the video frame resolution on a large time scale to achieve an efficient and fair user experience. Based on scalable video coding (SVC) technology and deep reinforcement learning (DRL) solutions, the present invention can output bandwidth and power allocation decisions at the time slot level by dynamically sensing the channel and traffic status; and simultaneously makes transmission selections for the enhancement layers of video frames at the frame level, thereby comprehensively optimizing the video frame transmission success rate, average resolution, and penalties such as latency and resolution degradation of the worst user.

[0058] As Figure 1 shown, the present invention provides a dual-time-scale wireless transmission method for video streaming, including:

[0059] Based on scalable video coding technology, each video frame in the video stream is divided into a base layer and multiple enhancement layers, and a correspondence relationship between the video frame resolution level transmitted to the user and the number of enhancement layers is established;

[0060] Construct a dual-time-scale decision-making mechanism: dynamically allocate bandwidth and power for users through scheduling and utility agents on the time slot-level time scale, and dynamically select the video frame resolution transmitted to the user through a resolution scaling agent on the video frame-level time scale;

[0061] The scheduling and utility agent and the resolution scaling agent are jointly trained using separate deep reinforcement learning network architectures;

[0062] Cross-layer state observation mechanism for the application layer and the transport layer: Construct the application layer state with the video frame period as the observation granularity and the transport layer state with the time slot as the observation granularity; When the transport layer state is input to the resolution scaling agent, it is converted into the average value of multiple time slot states within the target video frame period;

[0063] Two-level reward feedback mechanism: The reward function of the scheduling and utility agent includes an instantaneous resource efficiency metric and a cross-frame transmission progress penalty term, where the resources include the bandwidth and power; The resolution scaling agent reward function is based on the video frame transmission result, and calculates the weighted value of the video frame success rate, average resolution level, delay penalty, and resolution downgrade penalty of the worst user.

[0064] Specifically, as Figure 2 shown, there are users in the system, and the video stream of each user is encoded using SVC. Each video frame in the video stream contains a base layer and several enhancement layers. Let represent the bandwidth allocated to user n at the th time slot; represent the power allocated to user n at the th time slot. The video will select whether to transmit the next enhancement layer according to the current network condition and user-side feedback on a large time scale.

[0065] In one embodiment, the scheduling and utility agent includes a bandwidth decision network, a power decision network, a resource evaluation network, and a target resource evaluation network; The bandwidth decision network and the power decision network are jointly optimized under the guidance of the same evaluation network;

[0066] The resolution scaling agent includes a resolution decision network, a resolution evaluation network, and a target resolution evaluation network; The resolution decision network is optimized under the guidance of the resolution evaluation network.

[0067] This architecture design realizes the decoupled optimization of resource scheduling and resolution adjustment, and at the same time ensures decision coordination through shared state and reward design to optimize the video transmission success rate, average resolution, video frame delay, and resolution downgrade of the worst user.

[0068] In one embodiment, the construction of the application layer state with the video frame period as the observation granularity and the transport layer state with the time slot as the observation granularity specifically includes:

[0069] The application layer state includes the video frame delay sequence and the resolution level history Video frame success indicator and resolution decision action history ; The transmission layer state includes queue length network throughput channel attenuation sequence and resource allocation action history . Indicates the number of video frames that the historical video frame differs from the current video frame (video frame ).

[0070] Specifically, when the transmission layer state is input to the resolution scaling agent, it needs to be converted into the average value of the states of multiple time slots within the target video frame period to form a cross-layer state fusion feature.

[0071] In one embodiment, the cross-layer state observation mechanism for the application layer and the transmission layer further includes:

[0072] Before the application layer state and the transmission layer state are input to the scheduling and utility agent and the resolution scaling agent, a gated recurrent unit is used to process the most recent consecutive set of historical values of the application layer state and the transmission layer state to extract the corresponding temporal features; among them, the transmission layer state uses an observation sequence with a time slot as the granularity when the scheduling and utility agent makes a decision, and is averaged within the video frame period when the resolution scaling agent makes a decision to form a state average with a video frame as the granularity; therefore, the average value of the transmission layer state of the video frame in the resolution scaling agent and the transmission layer state of the time slot in the scheduling and utility agent satisfy the relationship:

[0073] ;

[0074] Among them, represents the set of time slot indices within the video frame , represents the number of time slots included in this video frame period, is the time slot index; the application layer state uses the state value with the same video frame time granularity when the scheduling and utility agent makes a decision and when the resolution scaling agent makes a decision.

[0075] In one embodiment, the two-level reward feedback mechanism specifically includes:

[0076] The reward in the scheduling and utility agent

[0077] ;

[0078] Among them, are the coefficients of the spectral efficiency reward, power efficiency reward, maximum remaining required rate user penalty, and video frame transmission success reward item or video frame transmission failure penalty item respectively; and are respectively the bandwidth ratio and power ratio allocated to user in time slot and are the spectral efficiency and power efficiency respectively, and are respectively the remaining data volume and remaining time of user in the current video frame, is a video frame transmission success indicator that takes effect at the end of each video frame,

[0079] The reward of the resolution scaling agent is:

[0080] ;

[0081] Among them, are the weight parameters of the video frame transmission success reward, video frame resolution reward, video frame resolution downgrade penalty, and video frame delay penalty respectively; is the number of video frames that user should transmit in one round of training; and are the number of video frame downgrade layers and video frame delay respectively, is the number of users in user set , is the total number of video frames in the video stream, represents the video frame of user

[0082] In one embodiment, the scheduling and utility agent and the resolution scaling agent are jointly trained using a separate deep reinforcement learning network architecture, specifically including:

[0083] Before each time slot, the scheduling and utility agent, based on state passes through the bandwidth decision network and power decision network of the scheduling and utility agent to obtain action , and after implementing the action, obtains the reward of the current time slot and the new state after implementing the action, and Stored in the buffer of the scheduling and utility agent, and network optimization update is performed after a batch size of time steps: Input the state into the target resource evaluation network, calculate the advantage function of the scheduling and utility agent based on the output and the reward, update the bandwidth decision network and the power decision network in the scheduling and utility agent based on the advantage function and the reward; Calculate the evaluation loss based on the advantage function and the output of the resource evaluation network, and update the resource evaluation network based on the evaluation loss; Regularly update the target resource evaluation network based on the resource evaluation network;

[0084] The resolution scaling agent is transmitting a video frame Before the period, based on the state The action is obtained through the resolution decision network of the resolution scaling agent After implementing the action, the video frame is obtained The reward And the new state after implementing the action Store in the buffer of the resolution scaling agent, and network optimization update is performed after a batch size of time steps: Input the state into the target resolution evaluation network, calculate the advantage function of the resolution scaling agent based on the output and the reward, update the resolution decision network in the resolution scaling agent based on the advantage function and the reward; Calculate the evaluation loss based on the advantage function and the output of the resolution evaluation network, and update the resolution evaluation network based on the evaluation loss; Regularly update the target resolution evaluation network based on the resolution evaluation network.

[0085] In one embodiment, dynamically allocating bandwidth and power for a user by the scheduling and utility agent on a time slot level time scale, and dynamically selecting the resolution of the video frame transmitted for the user by the resolution scaling agent on a video frame level time scale specifically includes:

[0086] Before the start of each time slot, the scheduling and utility agent receives the application layer state and the transport layer state processed by the time series, and the output actions include, under the premise of satisfying the maximum bandwidth constraint and the maximum power constraint, the bandwidth ratio allocated to the user in the time slot and the power ratio :

[0087] ;

[0088] ;

[0089] The user The allocated bandwidth and the power are respectively:

[0090] , ;

[0091] is the maximum available bandwidth for allocation, represents the maximum available power for allocation;

[0092] At the end of each video frame transmission, the resolution scaling agent receives the application layer state and transport layer state processed through time series, and the output action is the encoded value of the N-bit L-value digital code of all user resolution decision actions:

[0093] ;

[0094] wherein, represents for user selected from resolution levels of the video frame resolution, represents not transmitting the video frame, represents transmitting the base layer, represents the base layer and enhanced layers.

[0095] The scheduled and utility agent and resolution scaling agent that have completed training obtain resource allocation strategies and resolution decisions at small time scales (time slot level time scale) and large time scales (video frame level time scale), respectively.

[0096] In a preferred embodiment, the video streaming dual time scale control mechanism proposed by the present invention mainly consists of small time scale bandwidth and power resource allocation control and large time scale frame resolution adjustment. It is mainly divided into the following steps:

[0097] (1) Sense relevant information:

[0098] At the initial stage of system operation and at the beginning of each time slot, the base station senses parameters such as the cache queue status and historical frame transmission conditions of each user, receives channel state information (channel gain) from the users, and uploads this data to the edge agent. The edge agent receives and integrates the above information, records the real-time network environment conditions of each user, and uses it as the input for subsequent resource decisions.

[0099] (2) Small time scale bandwidth and power resource allocation:

[0100] At the beginning of each time slot, based on the aforementioned state information transmitted by the base station, the edge agent calls the trained deep reinforcement learning decision module to calculate the allocation of available bandwidth and power. Subsequently, the edge agent sends the allocation instructions to the base station, and the base station accordingly performs the operations of allocating bandwidth and power to each user, enabling each user to obtain the corresponding transmission resources in the current time slot.

[0101] (3) Large time-scale frame resolution adjustment:

[0102] At the end of several time slots or a video frame period, based on data such as the average channel gain of users, the status of the buffer queue, and the historical video frame transmission situation, the edge agent determines whether to enable a higher / lower resolution level for the next video frame. The decision result is uniformly sent by the agent at the cycle boundary, and the base station and users accordingly adjust the encoding and transmission settings of subsequent frames.

[0103] As Figure 3 shown, there are fast fluctuations at small time scales and channel changes at large time scales in the wireless channel. If an inappropriate resource allocation scheme is adopted, for example, Figure 3 a resource allocation scheme that remains unchanged all the time, it will lead to rapid fluctuations in the channel capacity. The channel capacity affects the transmission rate of the upper-layer video and thus affects the video bit rate. Therefore, the fluctuating channel capacity not only results in low resource efficiency but also causes fluctuations in the upper-layer video resolution, seriously affecting the user experience. Based on the dual time-scale decision scheme, a relatively stable channel capacity will be achieved, thereby improving the resource utilization efficiency and the smoothness of the upper-layer video transmission performance.

[0104] As Figure 4 shown, the application layer state and the transport layer state are respectively processed by a gated recurrent unit to extract time series features. The concatenated features are input into three decision networks, interact with the environment after obtaining actions, and store the updated required data in the replay buffer. Every certain number of time steps, samples are drawn from the buffer to update the network. Based on the target evaluation network and relevant data, the advantage function (Advantage Function, ADV) is obtained and the decision network is updated; in addition, the loss function of the evaluation network is obtained and the evaluation network is updated. The evaluation network will regularly update the parameters of the target evaluation network to better estimate the state quality. Figure 4 In it, GAE represents Generalized Advantage Estimation.

[0105] It should be understood that although the steps in the flowchart of the accompanying drawings of the specification are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings of the specification may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of the steps or stages in other steps or other steps.

[0106] Based on the description of the above method embodiments, the present invention also provides a system. The system may be a system that uses software (application), module, component, server, client, etc. described in the embodiments of this specification and combines necessary implementation hardware. Based on the same innovative concept, the systems in one or more embodiments provided by the embodiments of the present disclosure are as described in the following embodiments. Since the implementation solutions for the system to solve problems are similar to those of the method, the implementation of the specific system in the embodiments of this specification may refer to the implementation of the foregoing method, and the repeated parts will not be elaborated. As used hereinafter, the term "module" or "modular" is a combination of software and / or hardware that can implement a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0107] A video streaming double-time-scale wireless transmission system, comprising:

[0108] Video frame layering module: Based on scalable video coding technology, each video frame in the video stream is divided into a base layer and multiple enhancement layers, and the correspondence between the video frame resolution level transmitted to the user and the number of enhancement layers is established;

[0109] Double-time-scale decision module: Dynamically allocate bandwidth and power for the user through scheduling and utility agents at the time scale of time slots, and dynamically select the video frame resolution transmitted to the user through a resolution scaling agent at the time scale of video frames;

[0110] Training module: The scheduling and utility agents and the resolution scaling agent adopt a separate deep reinforcement learning network architecture for joint training;

[0111] Cross-layer state observation module: Construct an application layer state with a video frame period as the observation granularity and a transport layer state with a time slot as the observation granularity; when the transport layer state is input to the resolution scaling agent, it is converted into the average value of multiple time slot states within the target video frame period;

[0112] Two - level reward feedback module: The reward functions of the scheduling and utility agents include instantaneous resource efficiency metrics and cross - frame transmission progress penalty terms, where the resources include the bandwidth and power; the resolution scaling agent reward function calculates the weighted value of the video frame success rate, average resolution level, latency penalty, and resolution downgrade penalty of the worst - performing user based on the video frame transmission results.

[0113] In one embodiment, the present invention also provides a computer - readable storage medium including instructions, such as a memory including instructions, and the above - mentioned instructions can be executed by a processor to complete the above - mentioned method. The storage medium can be a computer - readable storage medium. For example, the computer - readable storage medium can be ROM, random access memory (RAM), CD - ROM, magnetic tape, floppy disk, and optical data storage devices, etc.

[0114] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above - mentioned exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non - restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.

[0115] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A dual-time-scale wireless transmission method for video streaming media, characterized in that including: Based on scalable video coding technology, each video frame in the video stream is divided into a base layer and multiple enhancement layers, and the corresponding relationship between the resolution level of the video frame transmitted to the user and the number of enhancement layers is established; Construct a dual-time-scale decision-making mechanism: on the time scale of time slots, the scheduling and utility agent dynamically allocates bandwidth and power for the user, and on the time scale of video frames, the resolution scaling agent dynamically selects the resolution of the video frame transmitted to the user; The scheduling and utility agent and the resolution scaling agent are jointly trained using a separate deep reinforcement learning network architecture; Cross-layer state observation mechanism for the application layer and the transport layer: construct the application layer state with the video frame period as the observation granularity and the transport layer state with the time slot as the observation granularity; when the transport layer state is input to the resolution scaling agent, it needs to be converted into the average value of multiple time slot states within the target video frame period; Dual-level reward feedback mechanism: the reward function of the scheduling and utility agent includes an instantaneous resource efficiency metric and a cross-frame transmission progress penalty term, and the resources include the bandwidth and power; the reward function of the resolution scaling agent is based on the video frame transmission result, and calculates the weighted value of the video frame success rate, average resolution level, delay penalty, and resolution degradation penalty of the worst user.

2. The dual-time-scale wireless transmission method for video streaming media according to claim 1, characterized in that The scheduling and utility agent includes a bandwidth decision network, a power decision network, a resource evaluation network, and a target resource evaluation network; the bandwidth decision network and the power decision network are jointly optimized under the guidance of the same resource evaluation network; The resolution scaling agent includes a resolution decision network, a resolution evaluation network, and a target resolution evaluation network; the resolution decision network is optimized under the guidance of the resolution evaluation network.

3. A dual-time-scale wireless transmission method for video streaming media according to claim 1, characterized in that The construction of the application layer state with the video frame period as the observation granularity and the transport layer state with the time slot as the observation granularity specifically includes: The application layer status includes a video frame delay sequence, a resolution level history, a video frame success indicator, and a resolution decision action history; the transport layer status includes a queue length, a network throughput, a channel attenuation sequence, and a resource allocation action history.

4. A dual-time-scale wireless transmission method for video streaming media according to claim 3, characterized in that The cross-layer state observation mechanism for the application layer and the transport layer further includes: Before the application layer state and the transport layer state are input into the scheduling and utility agent and the resolution scalability agent, the gated recurrent unit is used to process the most recent consecutive set of historical values of the application layer state and the transport layer state to extract the corresponding temporal features; among them, the transport layer state uses the observation sequence with the time slot as the granularity during the decision-making of the scheduling and utility agent, and performs the averaging process within the video frame period during the decision-making of the resolution scalability agent to form the state average value with the video frame as the granularity; therefore, the transport layer state average value of the video frame in the resolution scalability agent and the transport layer state of the time slot in the scheduling and utility agent satisfy the relationship: ​​ ; Among them, represents the set of time slot indices within a video frame, represents the number of time slots included in this video frame period, is the time slot index; the application layer state adopts the state value with the same video frame time granularity during scheduling and utility agent decision-making, as well as during resolution scalability agent decision-making.

5. A dual-time-scale wireless transmission method for video streaming media according to claim 4, characterized in that The dual-level reward feedback mechanism specifically includes: Reward in the Scheduling and Utility Agent is as follows: ; wherein, are the coefficients of spectral efficiency reward, power efficiency reward, maximum remaining required rate user penalty, and video frame transmission success reward item or video frame transmission failure penalty item respectively; and are respectively the bandwidth ratio and power ratio allocated to user in time slot , and are spectral efficiency and power efficiency respectively, and are respectively the remaining data volume and remaining time of user in the current video frame, is a video frame transmission success indicator that takes effect at the end of each video frame, is the user set; Reward of the Resolution Scaling Agent is as follows: ; Among them, are the weight parameters of the video frame transmission success reward, the video frame resolution reward, the video frame resolution degradation penalty, and the video frame delay penalty, respectively; In one round of training, the user should transmit the number of video frames; and are the number of video frame degradation levels and the video frame delay, respectively, is the number of users in the user set ; is the total number of video frames in the video stream, represents the user 's video frame resolution level.

6. A dual-time-scale wireless transmission method for video streaming media according to claim 5, characterized in that, The dynamic allocation of bandwidth and power for the user by the scheduling and utility agent on the time scale of time slots and the dynamic selection of the resolution of the video frame transmitted to the user by the resolution scaling agent on the time scale of video frames specifically includes: Before the start of each time slot, the scheduling and utility agents receive the application layer state and transport layer state processed by the time series, and the output actions include, subject to the maximum bandwidth constraint and the maximum power constraint, the bandwidth ratio allocated to the user and the power ratio : : ; ; User Allocated bandwidth and power are respectively as follows: , ; is the maximum bandwidth available for allocation, represents the maximum power available for allocation; At the end of each video frame transmission, the resolution scaling agent receives the application layer state and the transport layer state that have been processed in a time series, and outputs an action which is the encoded value of the N-bit L-value digital code for all user resolution decision actions: ; Among them, represents the video frame selected by the user from resolution levels means not to transmit the video frame, means to transmit the base layer, means the base layer and enhancement layers.

7. A video streaming double-time-scale wireless transmission method according to claim 6, characterized in that The joint training of the scheduling and utility agent and the resolution scaling agent using a separate deep reinforcement learning network architecture specifically includes: Before each time slot, the scheduling and utility agent makes a decision based on the state Through the bandwidth decision network and power decision network of the scheduling and utility agent, an action is obtained , and the reward for the current time slot is obtained after implementing the action and the new state after implementing the action , and is stored in the buffer of the scheduling and utility agent. After a batch size of time steps, network optimization and update are performed: the state is input into the target resource evaluation network, the advantage function of the scheduling and utility agent is calculated based on the output and the reward, and the bandwidth decision network and power decision network within the scheduling and utility agent are updated based on the advantage function and the reward; the evaluation loss is calculated based on the advantage function and the output of the resource evaluation network, and the resource evaluation network is updated based on the evaluation loss; the target resource evaluation network is periodically updated based on the resource evaluation network; The resolution scaling agent is transmitting a video frame Before the cycle, based on the state The action is obtained through the resolution decision network of the resolution scaling agent After implementing the action, the video frame is obtained Reward And the new state after implementing the action Will Stored in the buffer of the resolution scaling agent, and network optimization update is performed after a batch size of time steps: The state is input into the target resolution evaluation network, the advantage function of the resolution scaling agent is calculated based on the output and the reward, and the resolution decision network in the resolution scaling agent is updated based on the advantage function and the reward; The evaluation loss is calculated based on the advantage function and the output of the resolution evaluation network, and the resolution evaluation network is updated based on the evaluation loss; The target resolution evaluation network is periodically updated based on the resolution evaluation network.

8. A dual-time-scale wireless transmission system for video streaming media, characterized in that, including: Video frame layering module: Based on scalable video coding technology, each video frame in the video stream is divided into a base layer and multiple enhancement layers, and the corresponding relationship between the resolution level of the video frame transmitted to the user and the number of enhancement layers is established; Dual-time-scale decision-making module: On the time scale of time slots, the scheduling and utility agent dynamically allocates bandwidth and power for the user, and on the time scale of video frames, the resolution scaling agent dynamically selects the resolution of the video frame transmitted to the user; Training module: The scheduling and utility agent and the resolution scaling agent are jointly trained using a separate deep reinforcement learning network architecture; Cross-layer state observation module: construct the application layer state with the video frame period as the observation granularity and the transport layer state with the time slot as the observation granularity; when the transport layer state is input to the resolution scaling agent, it needs to be converted into the average value of multiple time slot states within the target video frame period. Two-level reward feedback module: The reward functions of the scheduling and utility agents include instantaneous resource efficiency metrics and cross-frame transmission progress penalty terms, where the resources include the bandwidth and power; the reward function of the resolution scaling agent is based on the video frame transmission results, and calculates the weighted value of the video frame success rate, average resolution level, delay penalty, and resolution downgrade penalty of the worst user.

9. A dual-time-scale wireless transmission system for video streaming according to claim 8, wherein The scheduling and utility agents include a bandwidth decision network, a power decision network, a resource evaluation network, and a target resource evaluation network; the bandwidth decision network and the power decision network are collaboratively optimized under the guidance of the same evaluation network. The resolution scaling agent includes a resolution decision network, a resolution evaluation network, and a target resolution evaluation network; the resolution decision network is optimized under the guidance of the resolution evaluation network.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Transmission control method of video-stream based on dual time scale

    WO2012079236A1

  • Antenna system precoding method and apparatus based on two time scales and deep learning

    WO2023035736A1